RAG and knowledge

RAG vs fine-tuning: which one your problem actually needs

RAG changes what the model can see. Fine-tuning changes how it behaves. Most projects need the first, some need the second, and a few need both.

Cover: RAG vs fine-tuning, which one your problem actually needs

Someone on the team says the chatbot “doesn’t know our product,” and within a week there is a proposal to fine-tune a model. It is an understandable instinct. Fine-tuning sounds like teaching, and teaching sounds like the fix for not knowing. It is usually the wrong tool for that particular complaint, and the reason is simple enough to fit in one sentence.

RAG changes what the model can see. Fine-tuning changes how the model behaves.

Almost every RAG vs fine-tuning decision falls out of that distinction, so it is worth being precise about both before comparing them.

What each one does to the model

Retrieval-augmented generation leaves the model untouched. When a question arrives, a retriever searches an index of your documents, picks the passages that look most relevant, and pastes them into the prompt next to the question. The model then answers with that material in front of it. The term comes from a 2020 paper by Patrick Lewis and colleagues, who described it as combining the model’s own “parametric” memory with a “non-parametric” one: a searchable index that sits outside the weights.

Fine-tuning goes the other way. You collect examples of inputs and the outputs you wish the model had produced, run a training job, and get back a model whose weights have moved. Nothing is looked up at question time. Whatever changed is now part of the model.

Diagram comparing a RAG pipeline, where a retriever pulls passages from editable documents into the prompt of an unchanged model, with fine-tuning, where training examples produce new model weights.
RAG leaves the model alone and changes its input. Fine-tuning changes the model.

Put side by side, the practical consequences are easy to read off. If a policy document changes on Tuesday, a RAG system reflects it as soon as the document is re-indexed. A fine-tuned model reflects it after someone builds a new training set and runs another job. If you need to show a user where an answer came from, RAG has the passage in hand; a fine-tuned model has nothing to point at.

What the research says about teaching a model facts

The intuition that you can train new facts into a model has been tested directly. In a December 2023 paper, Oded Ovadia and co-authors compared unsupervised fine-tuning against retrieval for injecting knowledge into LLMs across several knowledge-intensive tasks. Their summary is blunt: RAG “consistently outperforms” fine-tuning, both for knowledge the model had seen during pre-training and for entirely new knowledge.

The same paper found that models “struggle to learn new factual information through unsupervised fine-tuning,” and that showing the model many variations of the same fact helped. That second finding matters in practice. It means a fine-tuning run that sees each fact once, in one phrasing, is close to the worst case, and most internal document collections look exactly like that.

OpenAI’s own fine-tuning guide points the same way from the vendor side. The uses it lists are classification, nuanced translation, producing content in a specific format, and correcting failures to follow instructions. Adding knowledge is not on the list.

If the complaint is “it doesn’t know X,” reach for retrieval. If the complaint is “it knows, but it keeps answering in the wrong way,” that is where fine-tuning starts to earn its cost.

When fine-tuning is the right call

None of this makes fine-tuning a bad idea. It makes it a specific one. The cases where it clearly wins share a pattern: the model already has the knowledge it needs, and the problem is consistency.

  • A strict output format that prompting gets right most of the time but not every time, such as a JSON schema downstream code depends on.
  • Classification or routing at high volume, where a small tuned model can replace a large general one.
  • A house style for tone, length, or terminology that you would otherwise restate in a long system prompt on every request.
  • Shorter prompts at scale. Behaviour trained into the weights no longer has to be described in the prompt, which saves input tokens on every call.

The data requirement is smaller than many people expect. OpenAI’s guide sets the minimum at 10 examples, says it sees improvements from 50 to 100, and recommends starting with 50 well-made demonstrations and evaluating before collecting more. The hard part is rarely volume. It is writing examples that are actually the output you want, because the model will copy their flaws as faithfully as their strengths.

Cost used to be the other barrier, and parameter-efficient methods have lowered it a lot. LoRA, published by Edward Hu and colleagues in 2021, trains small low-rank matrices alongside frozen weights instead of updating the whole model. Against full fine-tuning of GPT-3 175B, the authors report 10,000 times fewer trainable parameters and a threefold drop in GPU memory. That is why a tuned adapter for an open model is now a realistic weekend project instead of a cluster booking.

When RAG is the right call

RAG is the default for anything that looks like “answer questions about our stuff”:

  • The source material changes, even occasionally.
  • Answers need citations, for users or for an auditor.
  • Different users are allowed to see different documents. Retrieval can filter by permission before anything reaches the prompt; a fine-tuned model cannot forget selectively.
  • You want to remove something later. Deleting a document from an index is an operation. Removing a fact from trained weights is a research problem.

RAG has its own failure modes, and they are worth naming because they are where projects actually stall. Retrieval can return the wrong passages, especially when documents are scanned, full of tables, or split into chunks that cut a sentence from its context. And a model can be handed the right passage and still answer from its own assumptions. These two problems need different fixes, which is why measuring them separately matters. We cover how in how to evaluate a RAG system.

A decision table you can argue with

Your situation Start with Why
Answers must come from internal documents RAG The model needs to see the text, not memorise it
Content changes weekly or daily RAG Re-index instead of retrain
Users need sources for each answer RAG Retrieved passages are the citation
Output format drifts despite a good prompt Fine-tuning Consistency is a behaviour problem
High-volume classification or extraction Fine-tuning A small tuned model can replace a large one
Long system prompt repeated on every call Fine-tuning, or prompt caching first Either removes or discounts the repeated tokens
Domain answers in a strict house style Both RAG for the facts, tuning for the shape

One row deserves a note. If your only reason to fine-tune is a long, repeated system prompt, check prompt caching before training anything. Providers now discount repeated prompt prefixes heavily, and that can remove most of the cost argument with no training run at all. The numbers are in our guide to cutting LLM costs.

Using both, in the right order

The combined setup is common and sensible: a model tuned for format and tone, fed retrieved passages for facts. The order in which you build it matters more than the architecture diagram suggests.

  1. Prompt first. Write the best instructions you can and a handful of examples. Keep a fixed set of test questions from day one.
  2. Add retrieval if the failures are about missing or stale information. Measure retrieval quality on its own, before you judge the answers.
  3. Fine-tune last, and only for failures that persist across many examples after retrieval is working. Train on outputs that include retrieved context, so the tuned model learns to use passages rather than ignore them.

Teams that start at step three tend to train a model on a problem that retrieval would have solved, then discover the tuned model still invents details about documents it never saw.

If you want to try RAG without building it

You do not need to write a pipeline to see whether retrieval solves your problem. A few open-source tools do the indexing, retrieval and citation for you, and trying one against twenty real questions will tell you more than a week of architecture discussion.

  • AnythingLLM is a desktop app: drop in documents, connect a local or hosted model, and chat with citations. It is the lightest way to test the idea on one machine, and we walk through a fully local setup in running RAG offline with Ollama.
  • RAGFlow is heavier and needs a real server, but its parser handles scans, tables and complex layouts, and it shows you the chunks so you can correct them.
  • Onyx is aimed at teams, with connectors that keep an index in sync with the tools your company already uses.

If the test shows retrieval is right but the build is more than you want to own, that is the kind of work our AI and automation engineering service takes on.

Comments

No comments yet — be the first to share what you think.

Leave a comment

Your email address stays private. Required fields are marked