AGENSPHERE/ JOURNAL
← JOURNAL
J-034DECISION4 MIN READ

RAG vs fine-tuning: which one your AI system actually needs

RAG vs fine-tuning is usually framed as a choice. It is not: RAG changes what the model reads, fine-tuning changes how it behaves. Use RAG for knowledge that changes or must be cited, fine-tuning for a stable format or task, and prompting before either.

IN SHORT
  • RAG (retrieval-augmented generation) puts relevant source material into the prompt at question time. Fine-tuning changes the model's weights by training it on examples.
  • Use RAG for knowledge: facts that change, documents that must be cited, and anything a user may not be allowed to see. Use fine-tuning for behaviour: a fixed output format, a tone, or a narrow task done at high volume.
  • Fine-tuning a model on your documents is a weak way to add knowledge. It is slow to update, cannot cite sources, cannot enforce permissions and still blurs rare facts.
  • Try in order: better prompts and examples, then RAG, then fine-tuning. Many production systems end up with both: a fine-tuned smaller model that reads retrieved context.

RAG vs fine-tuning is one of the first questions teams ask when a general model does not know enough about their business. It is usually framed as a choice between two ways of making the model smarter. That framing causes most of the mistakes, because the two techniques change different things.

01What is the difference between RAG and fine-tuning?

Retrieval-augmented generation (RAG) leaves the model alone. At question time, the system searches your sources, puts the most relevant passages into the prompt, and asks the model to answer from them. The full pipeline is in how RAG works, step by step.

Fine-tuning changes the model itself. You train it further on examples of inputs and the outputs you want, and its weights shift so that it produces that kind of output by default. How this fits with pretraining is in how LLMs are trained.

RAGFine-tuning
What changesWhat the model reads, per requestHow the model behaves, permanently
Best forKnowledge: facts, documents, policiesBehaviour: format, tone, a narrow task
UpdatingRe-index a document; live in minutesCollect examples, retrain, re-evaluate
CitationsYes, from the retrieved passagesNo; the source is lost in the weights
PermissionsEnforced in the search queryNot possible; everything is in the weights
Data neededYour documents, as they areHundreds to thousands of clean examples
Typical failureThe right passage was not retrievedConfident answers in the right format, with wrong facts

02When is RAG the right choice?

Choose RAG when the problem is that the model does not know something:

  • The knowledge changes. Prices, policies, product specs, tickets, contracts. Re-indexing a changed document is cheap; retraining a model for every change is not.
  • Answers must be traceable. RAG answers can point to the passage they came from, so a person can check them.
  • Different users may see different things. Permissions belong in the retrieval query, before anything reaches the model (see filter by permission before you rank). A fine-tuned model cannot forget what it learned for one user when another asks.
  • The corpus is large. No amount of training reliably stores thousands of rare facts; search finds them.

03When is fine-tuning the right choice?

Choose fine-tuning when the model knows enough but does not behave the way you need, consistently:

  • A strict output format that prompting gets right most of the time but not every time, at high volume.
  • A narrow, repeated task such as classification, extraction or routing, where a small fine-tuned model can match a large general one at a fraction of the cost per call (see cut LLM cost in the right order).
  • A house style or tone that is hard to describe in instructions but easy to show in examples.
  • Shorter prompts. Behaviour that is trained in does not need to be restated in every request, which saves tokens and latency.

Parameter-efficient methods such as LoRA make fine-tuning cheaper to run and easier to version, but they do not change what it is good for.

04Why not fine-tune the model on your documents?

It is the most common mistake in this decision. Training on documents teaches the model the style of your documents more reliably than their facts. The result is a model that sounds like your company and still invents the details, and it does so with more confidence than before. It cannot cite where an answer came from, it cannot respect who is allowed to see what, and every policy change means another training run.

If the goal is "the model should know our documents", the answer is retrieval.

05What should you try first?

Neither. Start with the prompt:

adaptation/order.py
def next_step(problem: str, evals: EvalReport) -> str:
    """The cheapest change that could fix the failure, in order."""
    if evals.failures_from("missing or stale knowledge"):
        return "rag"                      # retrieve it; do not train it in
    if evals.failures_from("instructions not followed"):
        return "prompt"                   # clearer instructions, examples, structured output
    if evals.failures_from("format or task drift at volume"):
        return "fine-tune"                # only once prompts have been tried and measured
    return "ship"                         # passing the eval set is the goal, not the technique

Clear instructions, a few good examples and structured outputs fix a surprising share of behaviour problems in hours, not weeks. Each step should be measured against the same evaluation set (see ship a prompt change with an eval), so you know the next, more expensive step is needed.

DECISION RECORD · DR-34LAYER L2
CONTEXTAn internal assistant answered questions about company policies. Answers were well written but often cited outdated rules, and the team proposed fine-tuning a model on the policy library.
OPTIONSFine-tune a model on the policy documents · RAG over the policy library with permission filters · RAG, plus a small fine-tuned model for the answer format
CHOSENRAG over the live policy library with permission filters and citations. Fine-tuning deferred until the eval set showed a format problem prompting could not fix.
REJECTEDFine-tuning on the documents would have frozen the policies at training time, removed citations and made every rule change a retraining job. The format layer was not needed once the prompt included two worked examples.

06Can you use RAG and fine-tuning together?

Yes, and mature systems often do. A common pattern is a smaller fine-tuned model that has learned the output format and the task, reading passages that retrieval supplies at question time. Retrieval brings the knowledge; fine-tuning makes the behaviour cheap and consistent. Each part is evaluated on its own: retrieval on whether the right passage was found, the model on whether it used it correctly.

RAG decides what the model reads. Fine-tuning decides how it writes. Most knowledge problems are reading problems.

07Where to go next

The terms used here are defined in the glossary. If your system is answering from the wrong sources, start with hybrid search and chunking by document structure. If it is answering in the wrong way, start with the eval set.

Building something like this?DESCRIBE A SYSTEM →