- Pretraining teaches a model to predict the next token over trillions of tokens of text. This is where it learns language, facts, code and reasoning patterns, and it is by far the most expensive stage.
- Supervised fine-tuning (instruction tuning) trains the pretrained model on examples of prompts with good responses, turning a text completer into an assistant.
- Preference tuning (RLHF, DPO and related methods) trains the model on comparisons between answers, so it favours responses people judge more helpful, honest and safe.
- Each stage leaves fingerprints: knowledge cutoffs come from pretraining, formatting habits from fine-tuning, and confident, agreeable answers partly from preference tuning.
This is part 4 of Foundations. Part 3 described the transformer's architecture. Here is how its billions of weights get their values, and why each stage matters to anyone building on top of a model.
01How does a network learn at all?
All neural network training follows the same loop:
- Give the model an input and let it make a prediction.
- Measure how wrong the prediction was with a loss function.
- Work out, for every weight, which direction would have reduced the loss (this is backpropagation, which computes gradients).
- Nudge every weight a tiny step in that direction (gradient descent).
- Repeat, billions of times.
No one tells the model what grammar is or what Paris is the capital of. It discovers whatever patterns reduce the loss.
02Stage 1: pretraining
Goal: predict the next token. Data: an enormous corpus, typically trillions of tokens, drawn from web pages, books, code, papers and other text, filtered and deduplicated.
For every position in every document, the model predicts the next token, and the loss is the cross-entropy between its predicted probabilities and the token that actually came next. Because the target is just the next token in existing text, no human labelling is needed. This is called self-supervised learning, and it is why the data can be so large.
def training_step(model, batch, optimizer):
inputs, targets = batch[:, :-1], batch[:, 1:] # targets are the inputs shifted by one token
logits = model(inputs) # (batch, seq, vocab) scores at every position
loss = cross_entropy(logits, targets) # how surprised was the model by the real next token?
loss.backward() # gradients for every weight
optimizer.step() # nudge the weights
optimizer.zero_grad()
return loss.item()To predict text well across that much variety, the model has to compress a great deal about the world: syntax, facts, styles, how code compiles, how arguments are structured. Pretraining a frontier model takes thousands of accelerators for weeks or months, which is why only a few organisations do it from scratch.
The result is a base model. It is a powerful text completer, but not an assistant. Ask it “What is the capital of France?” and it may continue with another quiz question, because that is a likely continuation of text that looks like a quiz.
03Stage 2: supervised fine-tuning (instruction tuning)
Goal: respond to instructions the way a helpful assistant would. Data: tens of thousands to millions of examples of prompts paired with high-quality responses, written or reviewed by people, plus increasingly synthetic examples checked for quality.
Training is the same next-token prediction, now on these curated conversations, usually with the loss computed only on the response. The model learns the format of a conversation, how to follow an instruction, when to use a list or code block, and how to decline.
This stage is where most of an assistant's style comes from.
04Stage 3: preference tuning
Goal: prefer answers people judge better. Data: comparisons. For one prompt, two or more responses are generated and people (or a model trained to imitate them) mark which is better.
Writing perfect answers is hard; judging which of two answers is better is much easier and scales further. Two well-known approaches:
- RLHF (reinforcement learning from human feedback): train a separate reward model to predict which answer people prefer, then use reinforcement learning to adjust the LLM to produce answers that score higher, while staying close to its fine-tuned behaviour.
- DPO (direct preference optimisation): skip the separate reward model and adjust the LLM directly so preferred answers become more likely than rejected ones.
More recent training also uses reinforcement learning with verifiable rewards: the model attempts problems whose answers can be checked automatically, such as maths results or code that must pass tests, and is rewarded for correct outcomes. This is a large part of how reasoning-focused models are trained.
05Why should builders care about how LLMs are trained?
Because each stage explains behaviour you will see in production:
| Behaviour | Comes from |
|---|---|
| Knows nothing after a certain date | Pretraining data cutoff |
| Fluent but sometimes wrong on niche facts | Knowledge compressed in weights during pretraining |
| Follows formats, uses markdown, refuses some requests | Supervised fine-tuning |
| Agreeable, confident, sometimes over-eager to please | Preference tuning rewards answers people like |
| Better at maths and code with “thinking” enabled | Reinforcement learning on verifiable tasks |
It also tells you what you can change. Very few teams pretrain. Some fine-tune open models for a narrow format or domain. Almost everyone gets most of the value from the cheaper levers: better prompts, better context through retrieval, and better systems around the model.
Pretraining gives a model knowledge, fine-tuning gives it manners, preference tuning gives it judgement about what people want. None of them gives it access to your data.
06When should you fine-tune instead of prompting?
Fine-tuning is worth it when you need a consistent format or style across very many calls, a smaller model to match a larger one on a narrow task, or behaviour that is hard to describe in a prompt but easy to show with examples. It is rarely the right way to add knowledge: facts that change belong in retrieval, covered in part 7.
07Where to go next
Training produces a model that outputs a probability for every possible next token. Part 5 explains how one token is actually chosen, and why the same prompt can give different answers.
Previous: Attention and the transformer. Next: Sampling: temperature, top-p and why answers vary.