AGENSPHERE/ JOURNAL
← JOURNAL
FOUNDATIONS · PART 7 OF 10
J-021GUIDE4 MIN READ

How RAG works, step by step

Retrieval-augmented generation (RAG) answers questions from your own documents by finding the relevant passages first and handing them to the model. Two pipelines, one for indexing and one for answering, and every quality problem lives in one of their steps.

IN SHORT
  • Retrieval-augmented generation (RAG) is a pattern: retrieve passages relevant to a question from your own data, put them in the prompt, and have the model answer from them.
  • It has two pipelines. Indexing (offline): load documents, split them into chunks, embed each chunk, store the vectors. Answering (per question): embed the question, search, assemble a prompt, generate an answer with citations.
  • RAG is how you give a model private or fresh knowledge without retraining it, and it is the main defence against hallucination on that knowledge.
  • Most RAG failures are retrieval failures: bad chunking, missed exact terms, missing permissions or stale indexes. Measure retrieval separately from the final answer.

This is part 7 of Foundations. Part 6 ended on the most effective fix for hallucination: give the model the facts. RAG is the standard way to do that with your own documents.

01What is RAG?

Retrieval-augmented generation is a pattern rather than a product. Before the model answers, the system retrieves relevant passages from a body of documents, augments the prompt with them, and the model generates an answer from that context.

It solves two problems at once. The model does not know your private data (it was never in its training set), and it does not know anything that changed after its training cutoff. RAG supplies both at question time, with no retraining.

02The two pipelines

A RAG system is two pipelines that share a store.

Indexing, which runs ahead of time and whenever documents change:

  1. Load documents from their sources: a wiki, a drive, a ticketing system, a database.
  2. Parse them into clean text with structure: headings, lists, tables.
  3. Chunk them into passages small enough to retrieve precisely.
  4. Embed each chunk with an embedding model, turning it into a vector (see part 2).
  5. Store the vectors with the text and metadata (source, section, permissions, last updated) in a vector database or search index.

Answering, which runs for every question:

  1. Embed the question with the same embedding model.
  2. Search for the chunks whose vectors are closest to the question, filtered to what this user may see.
  3. Re-rank the candidates and keep the best few.
  4. Assemble a prompt: instructions, the chosen chunks with IDs, and the question.
  5. Generate an answer, citing the chunk IDs it used.
rag/answer.py
def answer(question: str, user: User) -> Answer:
    q_vec = embed(question)
    candidates = index.search(q_vec, filter={"acl": user.principals}, top_k=30)
    top = rerank(question, candidates)[:5]
    sources = {f"S{i+1}": c for i, c in enumerate(top)}
    prompt = GROUNDED_PROMPT.format(
        sources="\n\n".join(f"[{sid}] {c.title}\n{c.text}" for sid, c in sources.items()),
        question=question,
    )
    text = llm.complete(prompt, temperature=0.2)
    return Answer(text=text, sources=sources)

That is a complete, if minimal, RAG system. Everything that makes it good in production is in the details of each step.

03Where does RAG go wrong?

Almost every RAG failure is a retrieval failure: the right passage never reached the model. Common causes, each with its own deeper note in this journal:

  • Bad chunks. A table split in half, a step separated from its heading. See chunk by document structure.
  • Missed exact terms. Vector search blurs product codes, error IDs and names. See hybrid search.
  • Permissions applied too late. Filtering after ranking leaks data and loses recall. See filter by permission before you rank.
  • Stale index. The source changed and the index did not. Re-index on change, not on a schedule nobody remembers.
  • Too much context. Sending twenty chunks when three answer the question dilutes attention and raises cost.

When retrieval works, generation failures are usually prompt issues: the model was not told to answer only from sources, or was not given a way to say it does not know.

04How do you measure a RAG system?

Measure the two halves separately, or you will not know which one to fix.

HalfQuestionMetric
RetrievalDid the answering passage make it into the top k?Recall at k
GenerationIs the answer correct and supported by the cited passages?Accuracy, faithfulness, citation validity

Build a set of real questions, each labelled with the passage that answers it, plus some questions your documents cannot answer. Run it on every change to chunking, embeddings, search or prompts.

05RAG or fine-tuning?

They solve different problems and are often used together.

  • Use RAG for knowledge: facts, documents, anything that changes or must be cited.
  • Use fine-tuning for behaviour: a format, a tone, a narrow task done consistently.

Fine-tuning a model on your documents is a weak way to add knowledge. It is expensive to update, cannot cite sources, and still blurs rare facts.

RAG does not make the model smarter. It makes sure the model is reading the right page.

06Where to go next

RAG decides what the model reads. The next step is letting the model decide what to do: call a function, query a system, take an action. Part 8 explains tool calling, the building block of every agent.

Previous: Why LLMs hallucinate. Next: How tool calling works.

Building something like this?DESCRIBE A SYSTEM →