AGENSPHERE/ JOURNAL
← JOURNAL
FOUNDATIONS · PART 6 OF 10
J-020GUIDE4 MIN READ

Why LLMs hallucinate, and what actually reduces it

LLMs hallucinate because a model generates the most plausible continuation, not the true one. When the facts are missing or blurry in its weights, plausible and true come apart. Grounding, verification and permission to say “I don't know” reduce it far more than prompts asking for accuracy.

IN SHORT
  • A hallucination is fluent, confident output that is false or unsupported, such as an invented citation, API or figure.
  • It happens because the model is trained to produce plausible text, its knowledge is compressed and approximate, and training often rewards answering over admitting uncertainty.
  • The most effective fixes are structural: give the model the facts in its context (grounding, often via retrieval), ask it to cite them, and check outputs against sources or rules.
  • Allow and reward “I don't know”: a system that can abstain is more trustworthy than one that always answers.

This is part 6 of Foundations. The first five parts covered how a model turns text into tokens, processes them and picks the next one. Those mechanics explain the failure everyone meets first: the model says something false, fluently and with total confidence.

01What is a hallucination?

A hallucination is output that sounds right but is false or not supported by any source: a paper that does not exist, a function that is not in the library, a policy clause nobody wrote, a statistic with no origin. The defining feature is not the error itself, people make errors too. It is that the false answer looks exactly like a true one.

02Why do LLMs hallucinate?

Several reasons stack on top of each other.

They optimise for plausible, not true. As part 1 showed, a model predicts the next token that is likely given the text so far. Most of the time, likely and true coincide, because true statements are common in training data. When the model lacks the fact, the most likely continuation is still something that looks like an answer.

Knowledge is compressed and blurry. Training packs trillions of tokens into a fixed number of weights. Common facts are stored robustly. Rare ones, such as a niche author's third paper or a version-specific API flag, are stored faintly or blended with similar facts. The model produces a confident-looking mix.

Training rewards answering. Fine-tuning examples mostly contain answers, and preference tuning (see part 4) tends to reward responses people like, which are usually confident and complete. Saying “I'm not sure” is underrepresented unless it is deliberately trained in.

There is no built-in fact check. The model does not look anything up while generating. Each token is chosen from the distribution, one at a time, and once a wrong token is chosen, the following tokens are generated to be consistent with it.

03What does not work very well?

  • Asking nicely. “Only state true facts” in a prompt helps a little; the model cannot follow an instruction to know things it does not know.
  • Lowering the temperature alone. It makes output more consistent, including consistently wrong. If the most likely answer is wrong, greedy decoding picks it every time.
  • A bigger model alone. Larger models hallucinate less on common knowledge, but still do on rare, recent or private facts, the ones businesses usually care about.

04What actually reduces hallucination?

1. Ground the answer in provided context. Put the relevant facts into the prompt and instruct the model to answer only from them. For your own data this is usually retrieval-augmented generation (RAG), walked through in part 7. The model is far better at reading and summarising a passage in front of it than at recalling a rare fact from its weights.

2. Require citations to the provided sources. Ask for the passage or document ID that supports each claim. Citations make unsupported claims visible, and you can check them automatically: does the cited passage exist, and does it contain the claimed fact?

3. Verify outputs in code. Many hallucinations are checkable. An order ID must exist. A function must exist in the library version you use. A number must match the source table. Validate at the boundary, as described in validate structured outputs.

answer/grounded.py
PROMPT = """Answer using only the sources below. Cite the source id in brackets after each claim.
If the sources do not contain the answer, reply exactly: I don't know based on the available documents.

Sources:
{sources}

Question: {question}"""

def check_citations(answer: str, sources: dict[str, str]) -> list[str]:
    problems = []
    for sid in re.findall(r"\[(S\d+)\]", answer):
        if sid not in sources:
            problems.append(f"cites {sid}, which was not provided")
    if "[" not in answer and "I don't know" not in answer:
        problems.append("no citations and no abstention")
    return problems

4. Give the model a way out. Explicitly allow “I don't know”, and treat it as a correct answer in your evals when the information is genuinely missing. A system that sometimes abstains is far more useful than one that always answers and is sometimes invented.

5. Use tools for facts that change. Prices, stock levels, account status and today's date should come from a tool call, not from the model's memory. Part 8 shows how.

6. Measure it. Keep a set of questions with known answers, including some whose answer is not in your documents. Track both accuracy and how often the system correctly says it does not know.

A model is a brilliant reader and an unreliable librarian. Hand it the right pages, and ask it to show which page it used.

05Is hallucination ever useful?

The same mechanism is what makes models creative: generating plausible text that has never existed is exactly what you want for brainstorming, naming or drafting fiction. The engineering task is not to eliminate it everywhere, but to make sure the system grounds and checks answers wherever truth matters.

06Where to go next

Grounding is the most important fix, and the standard way to ground a model in your own data is retrieval. Part 7 walks through a RAG system step by step.

Previous: Sampling: temperature and top-p. Next: How RAG works, step by step.

Building something like this?DESCRIBE A SYSTEM →