- RAG (retrieval-augmented generation) puts relevant source material into the prompt at question time. Fine-tuning changes the model's weights by training it on examples.
- Use RAG for knowledge: facts that change, documents that must be cited, and anything a user may not be allowed to see. Use fine-tuning for behaviour: a fixed output format, a tone, or a narrow task done at high volume.
- Fine-tuning a model on your documents is a weak way to add knowledge. It is slow to update, cannot cite sources, cannot enforce permissions and still blurs rare facts.
- Try in order: better prompts and examples, then RAG, then fine-tuning. Many production systems end up with both: a fine-tuned smaller model that reads retrieved context.
RAG vs fine-tuning is one of the first questions teams ask when a general model does not know enough about their business. It is usually framed as a choice between two ways of making the model smarter. That framing causes most of the mistakes, because the two techniques change different things.
01What is the difference between RAG and fine-tuning?
Retrieval-augmented generation (RAG) leaves the model alone. At question time, the system searches your sources, puts the most relevant passages into the prompt, and asks the model to answer from them. The full pipeline is in how RAG works, step by step.
Fine-tuning changes the model itself. You train it further on examples of inputs and the outputs you want, and its weights shift so that it produces that kind of output by default. How this fits with pretraining is in how LLMs are trained.
| RAG | Fine-tuning | |
|---|---|---|
| What changes | What the model reads, per request | How the model behaves, permanently |
| Best for | Knowledge: facts, documents, policies | Behaviour: format, tone, a narrow task |
| Updating | Re-index a document; live in minutes | Collect examples, retrain, re-evaluate |
| Citations | Yes, from the retrieved passages | No; the source is lost in the weights |
| Permissions | Enforced in the search query | Not possible; everything is in the weights |
| Data needed | Your documents, as they are | Hundreds to thousands of clean examples |
| Typical failure | The right passage was not retrieved | Confident answers in the right format, with wrong facts |
02When is RAG the right choice?
Choose RAG when the problem is that the model does not know something:
- The knowledge changes. Prices, policies, product specs, tickets, contracts. Re-indexing a changed document is cheap; retraining a model for every change is not.
- Answers must be traceable. RAG answers can point to the passage they came from, so a person can check them.
- Different users may see different things. Permissions belong in the retrieval query, before anything reaches the model (see filter by permission before you rank). A fine-tuned model cannot forget what it learned for one user when another asks.
- The corpus is large. No amount of training reliably stores thousands of rare facts; search finds them.
03When is fine-tuning the right choice?
Choose fine-tuning when the model knows enough but does not behave the way you need, consistently:
- A strict output format that prompting gets right most of the time but not every time, at high volume.
- A narrow, repeated task such as classification, extraction or routing, where a small fine-tuned model can match a large general one at a fraction of the cost per call (see cut LLM cost in the right order).
- A house style or tone that is hard to describe in instructions but easy to show in examples.
- Shorter prompts. Behaviour that is trained in does not need to be restated in every request, which saves tokens and latency.
Parameter-efficient methods such as LoRA make fine-tuning cheaper to run and easier to version, but they do not change what it is good for.
04Why not fine-tune the model on your documents?
It is the most common mistake in this decision. Training on documents teaches the model the style of your documents more reliably than their facts. The result is a model that sounds like your company and still invents the details, and it does so with more confidence than before. It cannot cite where an answer came from, it cannot respect who is allowed to see what, and every policy change means another training run.
If the goal is "the model should know our documents", the answer is retrieval.
05What should you try first?
Neither. Start with the prompt:
def next_step(problem: str, evals: EvalReport) -> str:
"""The cheapest change that could fix the failure, in order."""
if evals.failures_from("missing or stale knowledge"):
return "rag" # retrieve it; do not train it in
if evals.failures_from("instructions not followed"):
return "prompt" # clearer instructions, examples, structured output
if evals.failures_from("format or task drift at volume"):
return "fine-tune" # only once prompts have been tried and measured
return "ship" # passing the eval set is the goal, not the techniqueClear instructions, a few good examples and structured outputs fix a surprising share of behaviour problems in hours, not weeks. Each step should be measured against the same evaluation set (see ship a prompt change with an eval), so you know the next, more expensive step is needed.
06Can you use RAG and fine-tuning together?
Yes, and mature systems often do. A common pattern is a smaller fine-tuned model that has learned the output format and the task, reading passages that retrieval supplies at question time. Retrieval brings the knowledge; fine-tuning makes the behaviour cheap and consistent. Each part is evaluated on its own: retrieval on whether the right passage was found, the model on whether it used it correctly.
RAG decides what the model reads. Fine-tuning decides how it writes. Most knowledge problems are reading problems.
07Where to go next
The terms used here are defined in the glossary. If your system is answering from the wrong sources, start with hybrid search and chunking by document structure. If it is answering in the wrong way, start with the eval set.