- At each step the model produces a probability distribution over its whole vocabulary. Sampling is the rule that turns that distribution into one chosen token.
- Temperature reshapes the distribution: below 1 makes likely tokens more likely (more focused), above 1 flattens it (more varied). Near 0 approaches always picking the top token.
- Top-k keeps only the k most likely tokens; top-p (nucleus sampling) keeps the smallest set whose probabilities add up to p. Both cut off the unlikely tail before sampling.
- Low temperature suits extraction, classification and tool calls; higher settings suit brainstorming and creative writing. Even at temperature 0, outputs are not guaranteed identical across runs.
This is part 5 of Foundations. Part 4 explained how a model learns to predict the next token. This part covers the last step of every prediction: picking which token actually comes next.
01What does the model actually output?
Not a word. At each step, the final layer produces a score (a logit) for every token in the vocabulary, often more than 100,000 of them. A softmax turns those scores into probabilities that sum to 1.
After “The capital of France is”, the distribution might put most of the probability on “ Paris”, a little on “ a”, “ the” and “ located”, and tiny amounts on everything else. Sampling is the rule for choosing one token from that distribution.
02What is greedy decoding?
The simplest rule: always pick the most probable token. It is predictable and good for short factual answers. Over long outputs it tends to get repetitive or stuck in loops, because the single most likely continuation is often bland and self-reinforcing.
03What does LLM temperature do?
Temperature divides every logit by a number T before the softmax.
- T below 1 (say 0.2): differences between scores are magnified. The top token gets even more probability. Output becomes focused and consistent.
- T = 1: the model's distribution as trained.
- T above 1 (say 1.3): differences shrink, the distribution flattens, unlikely tokens get picked more often. Output becomes varied, and eventually incoherent.
- T near 0: effectively greedy.
import numpy as np
def sample(logits: np.ndarray, temperature=0.7, top_p=0.9, rng=np.random.default_rng()):
if temperature <= 0:
return int(np.argmax(logits)) # greedy
z = logits / temperature
probs = np.exp(z - z.max()); probs /= probs.sum() # softmax
order = np.argsort(probs)[::-1] # most likely first
cum = np.cumsum(probs[order])
keep = order[: np.searchsorted(cum, top_p) + 1] # nucleus: smallest set reaching top_p
p = probs[keep] / probs[keep].sum() # renormalise
return int(rng.choice(keep, p=p))04What are top-k and top-p?
Both remove the long tail of unlikely tokens before sampling, so a single strange low-probability choice cannot derail the output.
- Top-k: keep only the k most likely tokens (for example 40) and sample among them.
- Top-p, also called nucleus sampling: sort tokens by probability and keep the smallest set whose probabilities add up to p (for example 0.9). When the model is confident, that set might be two tokens; when it is unsure, it might be hundreds. That adaptivity is why top-p is the more common choice.
Temperature and top-p interact, so a common practice is to adjust one and leave the other at its default.
05Why do I get different answers to the same prompt?
Because sampling is random by design. With temperature above 0, each run draws different tokens, and one different early token sends the rest of the output down a different path.
Even at temperature 0, outputs are not guaranteed to be identical. Large models run on GPUs with parallel floating-point arithmetic, and the results can vary slightly with batch size and hardware. When two tokens are nearly tied, that tiny difference can flip the choice.
The practical lesson: never design a system that depends on a model producing byte-identical output twice. If you need repeatability, record what the model produced and replay the record, which is the idea behind the durable runtime in our KEEL teardown.
06What settings should I use?
| Task | Temperature | Why |
|---|---|---|
| Extraction, classification, tool selection | 0 to 0.3 | You want the most likely, consistent answer |
| Structured output (JSON) | 0 to 0.3 | Fewer format and value errors |
| Customer-facing replies | 0.3 to 0.7 | Natural wording without drifting |
| Brainstorming, creative writing | 0.8 to 1.1 | Variety is the point |
These are starting points. The right value for your system is the one that scores best on your eval set, as described in ship a prompt change with an eval.
The model gives you a distribution. Sampling is where you decide how much surprise you can afford.
07What else is a sampling setting?
A few other parameters work on the same distribution: max tokens caps output length, stop sequences end generation when a given string appears, frequency and presence penalties discourage repetition, and logit bias can push specific tokens up or down. Structured output modes go further and restrict sampling to tokens that keep the output valid against a schema.
08Where to go next
Sampling also helps explain a model's most famous failure. Part 6 looks at why models state false things fluently, and what actually reduces it.
Previous: How LLMs are trained. Next: Why LLMs hallucinate.