LLM and RAG engineering Guide

When to Fine-Tune an LLM in 2026: What It Fixes, What It Costs and Why the Market Moved

Fine-tuning was once the default answer to "how do we make the model ours." In 2026 OpenAI began closing its fine-tuning platform, and the evidence shows why: retrieval teaches facts better, and each new general model overtakes the last specialized one. Fine-tuning still wins on narrow, high-volume tasks. This guide sets out where, what it costs and how to run a project that pays.

For CTOs, heads of AI and product leaders deciding whether to fine-tune a language model, use retrieval, or switch models.

Published
Reviewed
Reading time
17 min

The short answer

Fine-tune a language model only when an evaluation shows a gap that prompting, retrieval or a different model cannot close, and the gap is about behavior: a format, a classification or extraction scheme, a tone, or a small model fast and cheap enough for high volume. Use retrieval for facts, and plan to retrain or retire the tuned model at every base-model release.

Key takeaways

  • Retrieval teaches facts better than fine-tuning. On 910 questions about events after the models' training cutoff, retrieval lifted Mistral 7B from 48.1% to 87.5%, and fine-tuning reached 50.4%.1
  • Specialized models keep losing to the next general model: medical fine-tunes beat their own base models in only 12.1% of comparisons, and GPT-4 with careful prompting passed Med-PaLM 2.23
  • Fine-tuning still wins on narrow, labeled tasks. A fine-tuned 0.5B model reached 0.83 F1 on relation extraction against 0.69 for GPT-5.4, and named deployments report large cost and latency savings.456
  • Managed fine-tuning retreated in 2026. OpenAI stops new training jobs for everyone on January 6, 2027, and Mistral, Databricks and Together have each pulled back.78910
  • Training is now the cheap part. Serving terms, retirement dates and retraining decide the cost, and a hosted fine-tune ends when its base model is retired.111213
  • Fine-tuning moves risk into the weights: ten harmful examples costing under $0.20 broke GPT-3.5 Turbo's guardrails, and tuning on a benign dataset raised its harmful-output rate from 5.5% to 31.8%.14

Fine-tuning means training an existing language model further on your own examples so that it behaves the way you need. For three years it was the default answer when a company wanted a model of its own. In 2026 the market moved. OpenAI began winding down its fine-tuning platform, other providers deprecated or narrowed theirs, and the strongest general models now match many tuned ones with a good prompt and the right documents. Fine-tuning still pays in specific cases, and those cases are easier to identify than they were. This guide covers what fine-tuning changes, where it loses and wins against the alternatives, what companies report, what it costs on each platform, which risks and rules come with it, and how to run a project that ends in a model worth keeping.

  • 9%of enterprise production models were fine-tuned in 2024, while 51% of enterprises used retrieval15
  • 12.1%of comparisons in which medical fine-tuned models beat the general model they were built on2
  • 31.8%harmful-output rate of GPT-3.5 Turbo after fine-tuning on a benign dataset, up from 5.5%14

What fine-tuning changes, and what it does not

The most useful finding for a buyer is that fine-tuning and retrieval solve different problems. Fine-tuning changes how a model behaves. Retrieval, usually called RAG, changes what it knows at the moment it answers by placing relevant documents in front of it. Microsoft researchers tested both on 910 questions about events after the models' training cutoff. Retrieval lifted Mistral 7B from 48.1% to 87.5%; fine-tuning on the same source text reached 50.4%, and for Llama 2 7B it cut accuracy from 35.3% to 21.9%.1 The authors conclude that models "struggle to learn new factual information through unsupervised fine-tuning."16 A study of less popular facts found that "RAG surpasses FT by a large margin."17

Exhibit 1Accuracy on 910 questions about events after the training cutoff, Mistral 7B
  • Base model48.1%
  • Fine-tuned on the source text50.4%
  • Fine-tuned on paraphrased text58.8%
  • Fine-tuned, then given retrieval81.0%
  • Base model with retrieval87.5%
Adding fine-tuning on top of retrieval lowered the score. For new facts, retrieval alone did best. Source: [1]

Fine-tuning on facts can also make a model worse. Google researchers found that examples carrying facts the model did not already know are learned "significantly slower," and once learned "they linearly increase the model's tendency to hallucinate."18 The practical rule follows: keep changing facts in retrieval, and build a fine-tuning set from examples whose facts the base model already handles. The two can add up when the target is behavior inside a domain. In a Microsoft agriculture study, fine-tuning added more than 6 percentage points of accuracy and retrieval added 5 more on top.19

Specialized models keep getting overtaken

The history of domain models is a history of being overtaken by the next general model. BloombergGPT, a 50 billion parameter model trained on a 363 billion token financial dataset,20 lost to five-shot GPT-4 on financial sentiment, with F1 scores of 0.86 against 0.51, and on conversational financial questions, 76.48 against 43.41.21 In medicine, GPT-4 with a prompting method called Medprompt outperformed the specialist Med-PaLM 2 "with an order of magnitude fewer calls" and passed 90% on the MedQA exam for the first time.3 A year later, o1-preview scored 96.0%, and few-shot prompting made it worse.22

A peer-reviewed audit of medical models built from general ones found that the medical versions beat their own base models in only 12.1% of three-shot comparisons, tied in 49.8% and were significantly worse in 38.2%.2 An adaptation tied to one base model has a shelf life measured in model generations, and a prompting recipe tuned to one generation can hurt the next.

Where fine-tuning still wins

The same financial study that embarrassed BloombergGPT shows where tuning keeps winning. On headline classification, a fine-tuned BERT model scored 95.36 against 86.00 for five-shot GPT-4.21 A 2024 study found fine-tuned small models "(still) significantly outperform" zero-shot generative models at text classification.23 In 2026, a fine-tuned Qwen2.5-0.5B reached a micro-F1 of 0.83 on relation extraction, against 0.69 for GPT-5.4 and 0.66 for Claude Sonnet 4.6 used zero-shot.4 Distillation compounds the effect: a 770 million parameter model trained on the reasoning of a 540 billion parameter model outperformed it on a benchmark using only 80% of the labeled data.24

The most quoted result in this area needs its conditions attached. LoRA Land fine-tuned 310 small models across 31 tasks and found they beat GPT-4 by 10 points on average, with 224 of 310 above GPT-4's score.25 GPT-4 was queried with "zero or single-shot, completion-style prompts," and the paper came from a company that sold fine-tuned model serving. Behavior also needs far less data than knowledge. LIMA tuned a 65 billion parameter model on 1,000 curated examples and was judged equal to or better than GPT-4 in 43% of comparisons.26 OpenAI's guidance is to start with 50 good examples and, "if 50 examples have no impact, rethink your task or prompt before adding training data."27

The gap you seeTry firstFine-tune when
Facts that change, or a large private document setRetrieval; the full documents in a long context if they fitRarely; at most a small tune for answer style on top of retrieval
Expert reasoning in a specialist fieldThe strongest current general model, well promptedYou have a grader that scores answers, and the next model release does not close the gap
A fixed format, schema, tone or house styleInstructions, examples and structured outputsThe prompt needs so many examples that cost or context runs out
High-volume classification, extraction or routingA frontier model, to set the bar and label dataVolume makes per-call cost or latency matter and labeled data exists
The same task, cheaper or fasterA smaller current model, promptedThe smaller model misses the bar and distillation closes it
The pattern in the evidence: retrieve for knowledge, tune for behavior, and test both against the next model release.

Long context is the other alternative. A Google DeepMind study found that when resourced, putting the whole text in the model's context "consistently outperforms RAG" on average, while retrieval costs far less, and that routing each question to one or the other kept quality close to long context at much lower cost.28 For a document set that fits in the window and is reused, cached long context removes much of the case for baking knowledge into weights.

What companies report

The production cases with numbers share one shape: a narrow, high-volume task moved from a frontier model to a small fine-tuned one, where the gain is mostly cost and speed. Checkr runs more than 1.5 million background checks; its engineer reported that a fine-tuned small model reached 97% accuracy against 88% for GPT-4, answered in about half a second, and cost about $800 a month against an estimated $7,000 to $12,000 with GPT-4.5 LinkedIn's domain-adapted Llama 3.1 8B models, trained on about 200 million tokens, were "75x and 6x cost effective" compared with GPT-4 and GPT-4o.6 Shopify runs 40 million multimodal inferences a day for its product catalog and cut median latency from 2 seconds to 500 milliseconds.29 A study that replaced a production GPT-4 feature with small open models measured a cost reduction of 5 to 29 times.30

Reinforcement fine-tuning, which trains a model against a grader that scores its answers, is the newer route, and its published results come from the seller. OpenAI reports that Ambience's medical coding model rose from 0.39 to 0.57, above a physician baseline of 0.45, and that Harvey's legal extraction F1 rose from 0.563 to 0.6765.31 The 2026 headline cases are larger projects on open weights the companies control. Harvey's Tenet, post-trained on the open Kimi K3 model, completes "almost twice as many held out tasks" as its base on Harvey's benchmark, at "roughly one-tenth the cost per cell" on one workload.32 Thomson Reuters spent $40 million training its own model from an open foundation.33 Both are company evaluations, and both are far from a typical project.

Adoption data shows fine-tuning is a minority practice. In Menlo Ventures' 2024 survey of 600 enterprise leaders, retrieval reached 51% adoption while "only 9% of production models" were fine-tuned.15 Andreessen Horowitz's 2025 survey of chief information officers reported "Fine-tuning viewed as less necessary as model capabilities improve," quoting one enterprise: "you just dump it into a long context and get almost equivalent results."34

Managed fine-tuning retreated in 2026

ProviderStatus, October 2026
OpenAIClosed to new organizations since May 7, 2026; closed to organizations without recent fine-tuned inference since July 2, 2026; no new jobs for anyone from January 6, 2027. Six fine-tuned model families, including fine-tuned o4-mini, shut down on October 23, 20267
AnthropicClaude 3 Haiku, the Claude model Amazon Bedrock offered for fine-tuning, reached end of life there on September 10, 20263536
MistralFine-tuning marked "deprecated" and "no longer actively supported"8
DatabricksFoundation Model Fine-tuning at end of life9
Together AITraining continues; "Serverless LoRA inference... has been discontinued," so tuned models need dedicated endpoints10
Google Vertex AIActive: supervised, preference and reinforcement tuning of Gemini models11
Amazon BedrockActive; once a base model turns Legacy, customization is restricted36
Microsoft AzureActive; fine-tuned GPT-4.1 deployments retire on October 14, 202713
Fireworks AIActive, and serves fine-tuned models "for the same price as base models"37
OpenAI's inference on existing fine-tuned models continues until each base model is deprecated.

OpenAI's documentation now opens its fine-tuning guides with "OpenAI is winding down the fine-tuning platform."27 The company has published no detailed reason. Its own model optimization guide says prompt engineering "may be all you need,"38 and the evidence above suggests why customers drifted away: for most tasks, a stronger general model with a better prompt and the right documents caught up with the tuned model, and a tuned model has to be rebuilt each time its base is retired.

What it costs

PlatformTraining priceServing a tuned model
Google Vertex AIPer 1,000 training tokens: Gemini 3.5 Flash $0.01; Gemini 2.5 Pro $0.025; Gemini 2.5 Flash $0.00511Tuned model endpoints listed at 1.5 times the base model price11
Microsoft AzureReinforcement fine-tuning of o4-mini at $100 per hour of training, capped at $5,000 per job12Base token price plus a $1.70 hourly hosting fee on Standard deployments12
Amazon BedrockPriced per model and token on its pricing page39Parameter-efficient tunes on demand or provisioned; full-rank tunes need Provisioned Throughput39
Together AILoRA from $0.34 per million tokens for small models to $7.00 for the largest40Dedicated endpoints only10
Fireworks AILoRA from $0.50 per million tokens up to 16B parameters37Same price as the base model37
Your own GPUsLambda lists an H100 at $4.29 and a B200 at $6.99 per GPU hour41Your serving stack; one GPU can serve many adapters42
Training costs tens of dollars for a typical behavior tune. The serving terms and the retraining cycle decide the bill.

At these rates, tuning on a million tokens for three epochs costs a few dollars on an open model and about $30 on Gemini 3.5 Flash. Training on your own hardware is in the same range: QLoRA showed in 2023 that a 65 billion parameter model can be fine-tuned "on a single 48GB GPU,"43 so a day of single-GPU training costs about $100 at Lambda's list price.41 The recurring bill is where platforms differ. Per-token billing at the base price makes a low-traffic tune almost free to keep. A 1.5 times markup scales with traffic. Azure's $1.70 hourly fee is about $1,200 a month per deployed model before any tokens.12 For teams running their own models, S-LoRA serves thousands of fine-tuned adapters with "up to 4 times" the throughput of naive serving, so one tune per task or per customer does not mean one GPU each.42 The larger costs are the data, the evaluation work and the retrain at every base-model change.

Risks that move into the weights

Fine-tuning can undo a provider's safety training. Researchers broke GPT-3.5 Turbo's guardrails by fine-tuning it on "only 10 such examples at a cost of less than $0.20," raising its harmful-output rate from 1.8% to 88.8%. The more important result for a company is the accidental one: one round of tuning on Alpaca, a widely used benign dataset, raised the rate from 5.5% to 31.8%.14 Narrow tuning can spread further than intended. Models fine-tuned to write insecure code without saying so began asserting "that humans should be enslaved by AI" on unrelated prompts, and a trigger could hide the behavior until it appeared.44

Training data can leak back out. A 2026 study showed that "small, proprietary SFT datasets can induce meaningful privacy leakage" of personal information in medical and legal settings.45 Replacing real records with generated data is no cure: after tuning on generated data, successful extractions of personal data rose by over 20% in one model family.46 Tuning also costs general ability. A study of continual fine-tuning found forgetting across models from 1 billion to 7 billion parameters, worsening as size grew,47 and LoRA, the common low-cost method, "substantially underperforms full finetuning" while better preserving skills outside the target task.48

A hosted fine-tune also has an end date. None of the closed platforms we reviewed documents a way to export tuned weights. Azure publishes retirement dates for fine-tuned deployments, such as October 14, 2027 for GPT-4.1,13 and Bedrock restricts customization once a base model turns Legacy.36 The lasting assets are the training data, the evaluation set and the recipe, so plan to keep and version them.

Licenses and rules

License or ruleWhat it saysWhat it means for a fine-tune
Llama 3.3 Community LicenseDistributed fine-tunes must include "Llama" at the beginning of the model name and display "Built with Llama"; above 700 million monthly users, a separate license is needed49Naming and attribution duties travel with any model you ship
Llama 4 acceptable use policyRights to the multimodal models are not granted to companies with a principal place of business in the European Union50EU companies cannot build on those models
Gemma 4Released under the Apache 2.0 license51Few duties beyond license notices; check each model card, since terms differ across families
Qwen2.5-72B licenseA license request above 100 million monthly users, and "Built with Qwen" on derivatives52Terms vary by model size within one family
Mistral Research LicenseModels and derivatives for "Research Purposes" only53No commercial use, including a fine-tune
Anthropic usage termsOutputs may train classifiers and extraction tools; general-purpose chatbots and open-ended generation models are prohibited54Distilling a frontier model is allowed only for narrow tools
EU AI Act, guidelines for general-purpose AIA company that modifies a model becomes its provider only if the modification uses more than a third of the original training compute55Almost no company fine-tune reaches this; obligations have applied since August 2, 202556
GDPR, EDPB Opinion 28/2024A model is anonymous only if extracting training data is "insignificant"; a model fine-tuned to mimic a person's voice cannot be anonymous57Treat a model tuned on personal data as holding that data
US copyright, Thomson Reuters v. ROSSThe Third Circuit held on September 29, 2026 that copying headnotes to train a competing AI tool was not fair use58Check the rights to every document in a training set
Most duties attach to the data in the weights and the license of the base model. Compute thresholds rarely apply.
Exhibit 2Three ways to adapt a language model

Prompt and retrieve

How we advise, by default

  • Facts stay current
  • Moves to a new model in days
  • Highest cost per call at volume
  • Evaluated on your own questions

Fine-tune a hosted model

Closed model, provider's platform

  • Quick to start
  • Ends when the base model retires
  • No weight export
  • Serving markups or hosting fees

Tune a small open model

When the evaluation shows a gap

  • Lowest cost and latency at volume
  • Weights stay with you
  • You run serving and safety tests
  • Retrained per base generation
Start with the first. Move a task to the third when it is narrow, high-volume and measured, and avoid the middle unless the platform's retirement dates suit you.

How to run a fine-tuning project that pays

Every provider gives the same first instruction: "Good evals first! Only invest in fine-tuning after setting up evals."27 A fine-tuning project is an evaluation project with a training step. The tuned model must beat three baselines: its own base model, the strongest current model with a good prompt and retrieval, and the next model release when it arrives. If it cannot beat the second, stop before training.

  1. Write the evaluation first Build a held-out set from real logged or expert-labeled cases, add a safety and general-ability check, and freeze it.
  2. Set the baselines Score the strongest current model with a careful prompt, retrieval or long context for knowledge, and a smaller model for cost.
  3. Name the gap Go ahead only if the remaining gap is about behavior, format, classification, or cost and latency at volume.
  4. Clear the ground Record the base model's license, the rights to each training source and the lawful basis for any personal data.
  5. Pilot with 50 examples Curate 50 to a few hundred examples, filter out facts the base model does not know, and remove personal data you do not need.
  6. Train small and test hard Start with LoRA on an open model, then rerun the task, safety and general-ability checks and an extraction test before release.
  7. Plan the retrain Version the data, evaluation and recipe, rerun the comparison at every major model release, and retire the tune when a prompted model matches it.

Questions before approving a fine-tuning project

  • Which evaluation shows a gap that the best prompted model with retrieval cannot close?
  • Is the gap about behavior and format, or about facts that belong in retrieval?
  • What happens to the tuned model when its base model is retired, and can we export the weights?
  • What does serving cost at our expected traffic, including hosting fees and markups?
  • Do the base model's license and our rights to the training data allow this use?
  • How will we test safety, data leakage and general ability before release?

The question has changed from "should we fine-tune GPT?" to "is this task narrow, high-volume and measurable enough to justify owning a small model?" When the answer is yes, the economics are better than ever, with training at tens of dollars, adapters that share hardware and permissive licenses. When it is no, a strong general model with retrieval will usually match the tuned model now and beat it after the next release. Either way, the evaluation set and the curated data are the assets that last. The weights can be rebuilt.

This is how we approach model customization in our AI development work: build the evaluation first, prove what prompting and retrieval achieve, and fine-tune a small open model only where the numbers show a gap worth owning, with the data, tests and recipe kept so it can be rebuilt on the next base model.

Questions leaders ask

What is fine-tuning an LLM?

Fine-tuning is further training of an existing language model on your own examples so it follows a format, a tone, a labeling scheme or a task more reliably. Common methods include supervised fine-tuning on example answers, preference tuning on pairs of better and worse answers, and reinforcement fine-tuning against a grader.

Fine-tuning vs RAG: which should we use?

Use retrieval for knowledge and fine-tuning for behavior. In a Microsoft study of new facts, retrieval lifted a 7B model from 48.1% to 87.5% while fine-tuning reached 50.4%. Fine-tuning helps when you need a fixed format, a classification or extraction scheme, a house style, or a smaller model that runs cheaply at volume.

Can we still fine-tune OpenAI models?

Only as an existing customer, and not for long. OpenAI's platform has been closed to new organizations since May 7, 2026, and no one can create new fine-tuning jobs from January 6, 2027. Existing fine-tuned models keep running until their base model is deprecated. Google Vertex AI, Amazon Bedrock, Azure and open-weight hosts still offer fine-tuning.

How much does it cost to fine-tune an LLM?

Training is usually cheap: a few dollars for a small open model, around $30 to tune Gemini 3.5 Flash on a million tokens for three epochs, or about $100 for a day on one rented H100. The larger costs are preparing data, building evaluations, serving markups or hosting fees, and retraining when the base model is retired.

How many examples do we need to fine-tune a model?

Fewer than most teams expect for behavior. OpenAI suggests starting with 50 good examples and rethinking the task if they show no improvement. The LIMA study used 1,000 curated examples. Teaching knowledge through examples works poorly at any size, so keep facts in retrieval.

Does fine-tuning reduce hallucinations?

Not when it teaches new facts. Google researchers found that once a model learns facts it did not already know through fine-tuning, its tendency to hallucinate rises linearly. Retrieval with citations is the better control. Fine-tuning can reduce format errors and off-task answers when the training examples use facts the model already knows.

Does the EU AI Act apply when we fine-tune a model?

Rarely as a model provider. Under the European Commission's guidelines, a company that modifies a general-purpose model becomes its provider only if the modification uses more than a third of the original model's training compute, which typical fine-tuning does not approach. The rules for the AI system you deploy still apply, as does GDPR for any personal data in the training set.

Sources

  1. Fine-tuning or retrieval? Comparing knowledge injection in LLMs (full paper)Ovadia et al., Microsoft
  2. Medical adaptation of large language and vision-language models: are we making progress?Jeong et al., EMNLP 2024
  3. Can generalist foundation models outcompete special-purpose tuning? Case study in medicineNori et al., Microsoft, 2023
  4. Sub-billion, super-frontier: small language models rival zero-shot frontier LLMs on general and literary relation extractionChristou and Tsoumakas, arXiv 2606.22606, 2026
  5. Checkr ditches GPT-4 for a smaller genAI model, streamlines background checksComputerworld, 2024
  6. How we built domain-adapted foundation GenAI models to power our platformLinkedIn Engineering, 2024
  7. DeprecationsOpenAI API documentation, read October 5, 2026
  8. Fine-tuning (deprecated)Mistral AI documentation, read October 5, 2026
  9. Create a training run using the Foundation Model Fine-tuning UI (EoL)Databricks documentation, read October 5, 2026
  10. LoRA vs. full fine-tuningTogether AI documentation, read October 5, 2026
  11. Agent Platform pricingGoogle Cloud, read October 5, 2026
  12. Fine-tuning cost managementMicrosoft Learn, read October 5, 2026
  13. Foundry Models lifecycle and support policyMicrosoft Learn, read October 5, 2026
  14. Fine-tuning aligned language models compromises safety, even when users do not intend to!Qi et al., ICLR 2024
  15. 2024: the state of generative AI in the enterpriseMenlo Ventures, November 2024
  16. Fine-tuning or retrieval? Comparing knowledge injection in LLMsOvadia et al., Microsoft, arXiv 2312.05934
  17. Fine tuning vs. retrieval augmented generation for less popular knowledgeSoudani et al., 2024
  18. Does fine-tuning LLMs on new knowledge encourage hallucinations?Gekhman et al., EMNLP 2024
  19. RAG vs fine-tuning: pipelines, tradeoffs, and a case study on agricultureBalaguer et al., Microsoft, 2024
  20. BloombergGPT: a large language model for financeWu et al., Bloomberg, 2023
  21. Are ChatGPT and GPT-4 general-purpose solvers for financial text analytics?Li et al., EMNLP 2023
  22. From Medprompt to o1: exploration of run-time strategies for medical challenge problems and beyondNori et al., Microsoft, 2024
  23. Fine-tuned small LLMs (still) significantly outperform zero-shot generative AI models in text classificationBucher and Martini, 2024
  24. Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizesHsieh et al., ACL 2023
  25. LoRA Land: 310 fine-tuned LLMs that rival GPT-4Zhao et al., Predibase, 2024
  26. LIMA: less is more for alignmentZhou et al., NeurIPS 2023
  27. Supervised fine-tuningOpenAI API documentation, read October 5, 2026
  28. Retrieval augmented generation or long-context LLMs? A comprehensive study and hybrid approachLi et al., Google DeepMind, EMNLP 2024
  29. Leveraging multimodal LLMs for Shopify's global catalogue: recap of expo talk at ICLR 2025Shopify Engineering
  30. Scaling down to scale up: a cost-benefit analysis of replacing OpenAI's LLM with open source SLMs in productionIrugalbandara et al., 2023
  31. Reinforcement fine-tuning use casesOpenAI API documentation, read October 5, 2026
  32. Post-training update: Harvey TenetHarvey, August 2026
  33. Thomson Reuters leverages its world-class data assets to launch its own frontier modelThomson Reuters, August 2026
  34. How 100 enterprise CIOs are building and buying gen AI in 2025Andreessen Horowitz, June 2025
  35. Fine-tuning for Anthropic's Claude 3 Haiku in Amazon Bedrock is now generally availableAWS, November 2024
  36. Model lifecycle (Legacy)Amazon Bedrock user guide, read October 5, 2026
  37. Fireworks AI pricingFireworks AI, read October 5, 2026
  38. Model optimizationOpenAI API documentation, read October 5, 2026
  39. Amazon Bedrock pricingAWS, read October 5, 2026
  40. Together AI pricingTogether AI, read October 5, 2026
  41. GPU cloud pricingLambda, read October 5, 2026
  42. S-LoRA: serving thousands of concurrent LoRA adaptersSheng et al., MLSys 2024
  43. QLoRA: efficient finetuning of quantized LLMsDettmers et al., NeurIPS 2023
  44. Emergent misalignment: narrow finetuning can produce broadly misaligned LLMsBetley et al., ICML 2025
  45. Reconstruction of personally identifiable information from proprietary data in supervised fine-tuned modelsFurukawa and Oprea, arXiv 2605.12264, 2026
  46. Generated data with fake privacy: hidden dangers of fine-tuning large language models on generated dataAkkus et al., USENIX Security 2025
  47. An empirical study of catastrophic forgetting in large language models during continual fine-tuningLuo et al., 2023
  48. LoRA learns less and forgets lessBiderman et al., TMLR 2024
  49. Llama 3.3 community license agreementMeta
  50. Llama 4 acceptable use policyMeta
  51. Gemma 4: expanding the Gemmaverse with Apache 2.0Google Open Source Blog, 2026
  52. Qwen2.5-72B-Instruct licenseAlibaba Cloud, via Hugging Face
  53. Mistral AI Research LicenseMistral AI
  54. Can I use my outputs to train an AI model?Anthropic Help Center, 2026
  55. Guidelines on the scope of the obligations for providers of general-purpose AI models under the AI ActEuropean Commission, July 2025
  56. Guidelines for providers of general-purpose AI modelsEuropean Commission
  57. Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI modelsEuropean Data Protection Board, December 2024
  58. Ease is not necessity: Third Circuit affirms no fair use in Thomson Reuters v. RossPatently-O, October 2026

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 5, 2026. No client data appears in our insights.

Read next

All insights
  • A person's question enters a lit retrieval hall at the centre of a plinth, which searches the company's document sources as that person, while one walled-off source has its lane barred in red. Only a few passages travel on to a small model, and every line of the answer carries a citation back to its source.

    LLM and RAG engineering Playbook

    Enterprise RAG in Production: Why Pilots Stall and What Fixes Them

    For CTOs deciding whether a RAG pilot that impressed in the demo can be trusted with real users, real permissions and real data.

    16 min read

  • Your own cases, stored in a golden-set archive, run through your system to a release gate whose board shows every segment against a threshold its owner signed in advance; one segment falls short and the gate holds, then the rerun clears the line and the release crosses a bridge to production.

    LLM and RAG engineering Guide

    LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship

    For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.

    16 min read

  • Three AI agents send their calls through one lit router, which passes most of them to a fleet of small models and only a hard one to a large frontier model. Each agent has its own budget gauge; one has spent to its cap and a red barrier stops it, while the other two keep working.

    Economics and buying Playbook

    AI Inference Cost: How to Govern LLM and Agent Spend

    For CFOs, CTOs and FinOps leads deciding how to forecast, allocate and cap the recurring cost of LLM applications and AI agents in production.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.