RAG vs Fine-Tuning an LLM: How to Choose

RAG vs Fine-Tuning an LLM: How to Choose

Retrieval-augmented generation (RAG) changes what a model reads at answer time: it finds relevant passages in your data and puts them in the prompt. Fine-tuning changes the model itself by training it further on your examples, which makes its behavior more consistent. Use RAG when the problem is missing or changing knowledge, and fine-tuning when the problem is format, tone or task skill. Build an evaluation set before either, because it tells you which problem you have.

What is the difference between RAG and fine-tuning?

RAG supplies knowledge from outside the model. Fine-tuning adjusts the model's weights. A study by Balaguer et al., posted January 16, 2024, puts it in one sentence: "RAG augments the prompt with the external data, while fine-Tuning incorporates the additional knowledge into the model itself."

OpenAI's guide to optimizing LLM accuracy (opened October 3, 2026) frames them as two separate levers:

  • Context optimization, which covers RAG, is needed when "the model lacks contextual knowledge because it wasn’t in its training set", "its knowledge is out of date", or "it requires knowledge of proprietary information."
  • LLM optimization, which covers fine-tuning, is needed when "the model is producing inconsistent results with incorrect formatting", "the tone or style of speech is not correct", or "the reasoning is not being followed consistently."

In the guide's words, "these are all levers that solve different things, and to optimize in the right direction you need to pull the right lever."

What does RAG change?

RAG changes the input and leaves the model untouched. The term comes from Lewis et al., posted May 22, 2020, who combined a language model's "parametric memory" with a searchable "non-parametric memory" and noted two open problems it addresses: "providing provenance for their decisions and updating their world knowledge".

Those two properties remain the reason to choose it. AWS's prescriptive guidance (opened October 3, 2026) lists them plainly: "RAG can incorporate the latest documents in a few minutes", and "a RAG model provides a reference to the information source."

RAG fails in its own ways. Barnett et al., posted January 11, 2024, documented seven failure points across three case studies: missing content, missed top-ranked documents, answer not in context, answer not extracted, wrong format, incorrect specificity and incomplete answers. The first three happen before the model writes a word.

Retrieval quality can be measured and improved. Anthropic reported on September 19, 2024 that adding context to each chunk and combining embeddings with keyword search "reduced the top-20-chunk retrieval failure rate by 49% (5.7% → 2.9%)", and that adding a reranker brought the reduction to "67% (5.7% → 1.9%)."

What does fine-tuning change?

Fine-tuning changes how the model behaves. OpenAI's guide gives the two usual reasons: "To improve model accuracy on a specific task" and "To improve model efficiency: Achieve the same accuracy for less tokens or by using a smaller model."

OpenAI's supervised fine-tuning documentation (opened October 3, 2026) lists what it suits: classification, nuanced translation, generating content in a specific format and correcting instruction-following failures. Google's tuning documentation, last updated October 2, 2026, adds an efficiency benefit: "Lower inference latency and cost due to shorter prompts".

Fine-tuning is a weak way to add facts. Three studies point the same way.

  • Ovadia et al., posted December 10, 2023, compared unsupervised fine-tuning with RAG on knowledge-intensive tasks: "RAG consistently outperforms it, both for existing knowledge encountered during training and entirely new knowledge."
  • Gekhman et al., posted May 9, 2024, found that "large language models struggle to acquire new factual knowledge through fine-tuning", and that as new facts are eventually learned, "they linearly increase the model's tendency to hallucinate."
  • AWS notes that "Fine-tuned models do not provide a reference to the source in their responses."

The training itself has become cheaper on open-weight models. The LoRA paper, posted June 17, 2021, showed that training small adapter matrices instead of all weights "can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times" compared with full fine-tuning of GPT-3.

What does each cost to build and maintain?

RAG costs more per request and less to keep current. Fine-tuning costs more up front and must be repeated when the data or the base model changes.

RAG Fine-tuning
What changes The prompt: retrieved passages are added The model weights
Up-front work Ingestion, chunking, an index, retrieval tuning A labeled dataset, training runs, a held-out test set
Data needed Your documents as they are Example inputs with correct outputs. OpenAI: minimum 10, start with 50. Google: "think 100 examples or more"
Updating knowledge Re-index. AWS: "in a few minutes" Retrain. AWS: "a few hours to days, depending on the size of the model"
Cost per request Higher: every request carries retrieved text Can be lower: shorter prompts or a smaller model
Sources in answers Yes, passages can be cited No
Ongoing upkeep Index freshness, access rules, retrieval evaluation Retraining when requirements or the base model change

Two maintenance facts deserve attention before a commitment.

A hosted fine-tune depends on the vendor's roadmap. OpenAI's deprecations page (opened October 3, 2026) records that "On May 7th, 2026, we notified developers using OpenAI’s self-serve fine-tuning platform of updates to availability." From January 6, 2027, "Active existing customers will no longer be able to create new fine-tuning jobs", and inference on fine-tuned models ends when the underlying base model is deprecated. An adapter trained on an open-weight model that you host does not carry that dependency.

A small knowledge base may need neither technique. Anthropic's guidance from September 2024: "If your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt".

When should you use RAG, fine-tuning or both?

Match the fix to the failure you observe in testing. The table below maps common symptoms to the cause and the first thing to try.

What you observe Likely cause First thing to try
The model does not know your products, policies or customers Missing knowledge RAG, or the full documents in the prompt if they are small
Answers were right last quarter and are wrong now Stale knowledge RAG with a re-indexing schedule
Users need to see where an answer came from No provenance RAG with cited passages
The right document exists but the answer ignores it Retrieval failure Fix chunking, add keyword search and reranking
The right passage was retrieved but the answer is still wrong Generation failure Better instructions, then fine-tune on examples that include retrieved context
Output format, structure or tone varies between runs Inconsistent behavior Tighter prompt and examples, then fine-tune
Prompts are long and the volume is high Cost per request Fine-tune a smaller model to replace instructions and examples
Nobody can say why answers are wrong No evaluation Build a test set before changing anything else

When should you combine RAG and fine-tuning?

Combine them when testing shows both a knowledge problem and a behavior problem. OpenAI's guide says so directly: "These techniques stack on top of each other - if your early evals show issues with both context and behavior, then it’s likely you may end up with fine-tuning + RAG in your production solution."

The measured effect can add up. In the agriculture case study by Balaguer et al., "We see an accuracy increase of over 6 p.p. when fine-tuning the model and this is cumulative with RAG, which increases accuracy by 5 p.p. further."

The combination works when the model is trained on the same kind of input it will see in production. OpenAI's advice: "if you have a RAG application, fine-tune the model with RAG examples in it". The RAFT method by Zhang et al., posted March 15, 2024, builds on that idea: "we train the model to ignore those documents that don't help in answering the question".

Why is fine-tuning often the wrong first move?

Because fine-tuning does not supply missing context, and its effect cannot be judged without a test set. A team that fine-tunes to add company knowledge is using the method the studies above found weaker for that purpose.

The providers who sell fine-tuning say to do other things first.

  • OpenAI's fine-tuning documentation: "Good evals first! Only invest in fine-tuning after setting up evals."
  • OpenAI's accuracy guide: "many of our largest customer deployments at OpenAI were done using only prompt engineering and RAG."
  • Google's tuning documentation: "We recommend starting with prompting to find the optimal prompt. Then, move on to fine-tuning (if required) to further boost performances or fix recurrent errors."

A practical order of work:

  1. Write the evaluation set. OpenAI's threshold for a useful baseline is "a set of 20+ questions and answers" plus a hypothesis for each failure.
  2. Sort the failures. For each wrong answer, check whether the right passage was retrieved. If it was not, the problem is retrieval. If it was, the problem is generation.
  3. Fix retrieval problems in the retrieval layer. Chunking, hybrid search and reranking address them. Training does not.
  4. Fix behavior with the prompt first, using instructions and examples.
  5. Fine-tune when behavior problems remain or the prompt has become expensive. OpenAI's rule of thumb: "If 50 examples have no impact, rethink your task or prompt before adding training data."

What is agentic RAG?

Agentic RAG lets the model control retrieval: it plans searches, runs several, judges the results and searches again when they fall short. Standard RAG runs one fixed search per question.

The survey by Singh et al., first posted January 15, 2025 and revised April 1, 2026, describes the motivation: traditional RAG systems "are constrained by static workflows and lack the adaptability required for multi-step reasoning and complex task management", and agentic RAG addresses this "by embedding autonomous AI agents into the RAG pipeline."

A production example is the agentic retrieval feature in Azure AI Search. Microsoft's documentation, updated September 16, 2026, describes "a multi-query pipeline designed for complex questions" that can "break down a complex query into smaller, focused subqueries" and that "Runs subqueries in parallel."

It suits questions with several parts or several data sources, at the price of more model calls per question. Anthropic measured on June 13, 2025 that "agents typically use about 4× more tokens than chat interactions". Start with standard RAG and add agentic retrieval when the evaluation set shows multi-part questions failing.

Key takeaways

  • RAG changes the prompt and suits private, new or changing knowledge and answers that cite sources.
  • Fine-tuning changes the weights and suits consistent format, tone and task skill, and shorter prompts at high volume.
  • Research finds fine-tuning a poor way to add facts: RAG "consistently outperforms it", and newly learned facts raise hallucination.
  • Both can be combined. One study measured over 6 percentage points from fine-tuning and a further 5 from RAG.
  • A hosted fine-tune depends on the provider. OpenAI ends new self-serve fine-tuning jobs on January 6, 2027.
  • Build the evaluation set first. It shows which lever to pull.

Frequently asked questions

Is RAG cheaper than fine-tuning?

RAG is usually cheaper to start and to keep current, because documents can be re-indexed in minutes with no training. It costs more per request, since every prompt carries retrieved text. Fine-tuning costs more up front and can lower the cost per request by shortening prompts or allowing a smaller model.

Can fine-tuning replace RAG?

For knowledge, no. Ovadia et al. found RAG ahead of unsupervised fine-tuning for both existing and new knowledge, and a fine-tuned model cannot cite its sources. Fine-tuning replaces long instructions and examples in the prompt. It does not replace a searchable knowledge base.

How much data do you need to fine-tune an LLM?

Tens to hundreds of examples to start. OpenAI sets a minimum of 10, recommends starting with 50 well-crafted ones, and reports improvements "from fine-tuning on 50–100 examples". Google suggests 100 or more. Both stress quality and similarity to production inputs over volume.

Does fine-tuning reduce hallucinations?

Not reliably. Gekhman et al. found that fine-tuning examples containing new facts "linearly increase the model's tendency to hallucinate" once learned, and AWS lists "an increased risk of hallucination" as a disadvantage of fine-tuned models for question answering. Grounding answers in retrieved passages is the more direct control.

Do you still need RAG with a large context window?

For a small, stable knowledge base, often not. Anthropic's guidance is that under 200,000 tokens, about 500 pages, the whole knowledge base can go in the prompt. Larger or frequently changing collections still need retrieval, both for cost and for accuracy.

Easital Technologies Ltd. builds RAG systems, covering ingestion, retrieval, evaluation and upkeep, and has trained and fine-tuned models, including on local and on-premise setups. Related pages: RAG development, LLM development, LLM cost optimization and private LLM deployment.

Sources

All sources were opened and checked on October 3, 2026.

  1. Balaguer et al., "RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture", arXiv 2401.08406, January 16, 2024. https://arxiv.org/abs/2401.08406
  2. OpenAI, "Optimizing LLM Accuracy", API documentation, undated, opened October 3, 2026. https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy
  3. Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv 2005.11401, May 22, 2020. https://arxiv.org/abs/2005.11401
  4. Amazon Web Services, "Comparing Retrieval Augmented Generation and fine-tuning", AWS Prescriptive Guidance, undated, opened October 3, 2026. https://docs.aws.amazon.com/prescriptive-guidance/latest/retrieval-augmented-generation-options/rag-vs-fine-tuning.html
  5. Barnett, Kurniawan, Thudumu, Brannelly, Abdelrazek, "Seven Failure Points When Engineering a Retrieval Augmented Generation System", arXiv 2401.05856, January 11, 2024. https://arxiv.org/abs/2401.05856
  6. Anthropic, "Contextual Retrieval in AI Systems", September 19, 2024. https://www.anthropic.com/news/contextual-retrieval
  7. OpenAI, "Supervised fine-tuning", API documentation, undated, opened October 3, 2026. https://developers.openai.com/api/docs/guides/supervised-fine-tuning
  8. Google Cloud, "Introduction to tuning", Gemini Enterprise Agent Platform documentation, last updated October 2, 2026. https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning
  9. Ovadia, Brief, Mishaeli, Elisha, "Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs", arXiv 2312.05934, December 10, 2023. https://arxiv.org/abs/2312.05934
  10. Gekhman, Yona, Aharoni, Eyal, Feder, Reichart, Herzig, "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?", arXiv 2405.05904, May 9, 2024. https://arxiv.org/abs/2405.05904
  11. Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", arXiv 2106.09685, June 17, 2021. https://arxiv.org/abs/2106.09685
  12. OpenAI, "Deprecations", API documentation, entry dated May 7, 2026. https://developers.openai.com/api/docs/deprecations
  13. Zhang, Patil, Jain, Shen, Zaharia, Stoica, Gonzalez, "RAFT: Adapting Language Model to Domain Specific RAG", arXiv 2403.10131, March 15, 2024. https://arxiv.org/abs/2403.10131
  14. Singh, Ehtesham, Kumar, Khoei, Vasilakos, "Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG", arXiv 2501.09136, January 15, 2025 (v4 April 1, 2026). https://arxiv.org/abs/2501.09136
  15. Microsoft, "Agentic Retrieval Overview", Azure AI Search documentation, updated September 16, 2026. https://learn.microsoft.com/en-us/azure/search/agentic-retrieval-overview
  16. Anthropic, "How we built our multi-agent research system", June 13, 2025. https://www.anthropic.com/engineering/multi-agent-research-system

Further reading

Work with Easital on a RAG system

Send a short description of the product or the problem you want solved. We reply by email with questions and a proposed next step.

Discuss your project