- Self-attention lets each token build its new representation from a weighted mix of the tokens before it, with weights computed from how relevant each one is.
- Each token produces three vectors: a query (what am I looking for?), a key (what do I contain?) and a value (what do I pass on?). Relevance is the dot product of a query with each key.
- A transformer layer is attention followed by a feed-forward network, with residual connections and normalisation. LLMs stack dozens of these layers.
- In a language model the attention is causal: a token can only look at earlier tokens, which is what makes next-token prediction possible.
This is part 3 of Foundations. In part 2 every token became a vector. Now we look at what the transformer does with those vectors, and in particular at attention, the idea from the 2017 paper “Attention Is All You Need” that made modern LLMs possible.
01What problem does attention solve?
The meaning of a word depends on the words around it. In “the bank raised its rates”, bank is a financial institution; in “we sat on the river bank”, it is not. its refers back to bank, several words earlier.
Before transformers, models read text one token at a time and squeezed everything seen so far into a single running state, which made long-range connections hard to keep. Attention takes a different approach: every token can look directly at every earlier token and decide how much each one matters.
02How does attention work?
For each token, the layer computes three vectors from its current representation, each by multiplying with a learned matrix:
- Query (Q): what this token is looking for.
- Key (K): what this token offers to others.
- Value (V): the information this token passes on if someone attends to it.
Then, for the token at position i:
- Compare its query with the key of every token up to i, using a dot product. A high score means “this one is relevant to me”.
- Scale the scores and pass them through a softmax, so they become weights between 0 and 1 that add up to 1.
- Take a weighted sum of the value vectors using those weights. That sum is the token's new, context-aware representation.
In matrix form this is the formula you will see everywhere: softmax(Q·Kᵀ / √d) · V, where d is the size of the key vectors. Dividing by √d keeps the scores in a range where softmax still gives useful gradients.
import numpy as np
def causal_self_attention(x, Wq, Wk, Wv):
"""x: (seq_len, d_model). Returns one context-aware vector per position."""
Q, K, V = x @ Wq, x @ Wk, x @ Wv
d = K.shape[-1]
scores = Q @ K.T / np.sqrt(d) # (seq_len, seq_len) relevance scores
mask = np.triu(np.ones_like(scores), k=1).astype(bool)
scores[mask] = -np.inf # causal: no looking at future tokens
weights = np.exp(scores - scores.max(axis=-1, keepdims=True))
weights /= weights.sum(axis=-1, keepdims=True) # softmax per row
return weights @ VThis is a toy version with real logic: about a dozen lines that are the heart of every LLM. Production code is the same idea, made fast on GPUs.
03Why “causal”?
A language model is trained to predict the next token, so during training it must not see the answer. The mask in the code sets the score for every future position to minus infinity, which becomes a weight of zero after softmax. Each token can only attend to itself and earlier tokens. That is what “causal” or “masked” self-attention means.
04What is multi-head attention?
One set of Q, K and V matrices learns one kind of relationship. Real models run many attention heads in parallel, each with its own matrices, then combine their outputs. Different heads tend to specialise: one may track which noun a pronoun refers to, another nearby syntax, another long-range topic. Nobody programs these roles; they emerge from training.
05What is in a full transformer layer?
Attention is one half. A transformer layer (also called a block) has:
- Multi-head self-attention: tokens exchange information.
- A feed-forward network: a small two-layer neural network applied to each token on its own. This is where much of the model's stored knowledge is thought to live.
- Residual connections: each sub-layer's output is added to its input rather than replacing it, which lets information and gradients flow through very deep stacks.
- Normalisation: keeps the numbers in a stable range between layers.
An LLM stacks this block dozens of times, sometimes more than a hundred. Early layers tend to capture local, surface patterns; later layers more abstract relationships. The output of the last layer feeds the prediction step from part 1.
Attention lets every token ask every earlier token one question: how much do you matter to me? A transformer asks it dozens of times, in parallel, at every layer.
06Why does attention get expensive with long inputs?
Every token compares itself with every earlier token, so the work grows with the square of the sequence length. Double the context and the attention work roughly quadruples. This is a big reason long context windows cost more and are slower, and why there is so much engineering around efficient attention and caching.
07Where to go next
We now have the whole architecture: tokens, embeddings, a stack of attention and feed-forward layers, and a prediction. But a freshly built transformer outputs noise. Part 4 explains how training turns it into something useful.
Previous: Tokens and embeddings. Next: How LLMs are trained.