- A large language model (LLM) is a neural network trained to predict the next token (a word or piece of a word) given the text so far.
- Generating a reply means running that prediction in a loop: predict a token, append it to the text, predict again, until a stop condition.
- Inside, text becomes token IDs, IDs become vectors (embeddings), a stack of transformer layers refines those vectors using attention, and a final layer turns the last vector into a probability for every token in the vocabulary.
- The model has no database of facts and no memory between requests. What it knows is stored in its weights from training, and what it sees is only the text in the current request.
This is part 1 of Foundations, a series that explains how large language models and AI agents work, from the first token to a production system. Each part builds on the one before, and each one is written for someone new to the field.
If you remember one sentence from this part, make it this: a large language model predicts the next token, over and over. Everything else is detail about how it does that well.
01What is a large language model?
A large language model (LLM) is a neural network, a very large mathematical function with billions of adjustable numbers called parameters or weights. It was trained on a huge amount of text to do one task: look at a sequence of text and predict what comes next.
“Large” refers to the number of parameters and the amount of training data. “Language model” is an older term for any system that assigns probabilities to sequences of words. The modern versions are built on an architecture called the transformer, introduced in 2017, which we unpack in part 3.
02What happens when you send a prompt?
Follow one request: you type “The capital of France is” and press enter.
1. Tokenisation. Your text is split into tokens: whole words, pieces of words, or punctuation. Each token maps to an integer ID from a fixed vocabulary of tens of thousands to a few hundred thousand entries. “The capital of France is” becomes something like five IDs.
2. Embedding. Each ID is looked up in a table that turns it into a vector: a list of a few thousand numbers. These vectors are how the model represents meaning. Information about each token's position is added too, because order matters. Part 2 covers both steps.
3. The transformer layers. The vectors pass through a stack of identical layers, often dozens of them. In each layer, two things happen. Attention lets every token look at the tokens before it and pull in relevant information (so “is” can take context from “capital” and “France”). A feed-forward network then transforms each token's vector on its own. Each layer refines the representation a little more.
4. The prediction. The final vector for the last position is multiplied by a matrix that produces one score, called a logit, for every token in the vocabulary. A softmax function turns those scores into probabilities that add up to 1. “ Paris” gets a high probability; “ banana” gets a tiny one.
5. Sampling. One token is chosen from that distribution. Usually the likely ones win, but how strictly depends on settings like temperature, covered in part 5.
6. Repeat. The chosen token is appended to the input, and the whole prediction runs again for the next position. This continues until the model produces a special end token or hits a length limit.
def generate(model, tokenizer, prompt, max_new_tokens=50):
ids = tokenizer.encode(prompt) # text -> token IDs
for _ in range(max_new_tokens):
probs = model.next_token_probs(ids) # one forward pass -> distribution over vocabulary
next_id = sample(probs, temperature=0.7) # pick one token
if next_id == tokenizer.eos_id: # model says it is done
break
ids.append(next_id) # feed it back in
return tokenizer.decode(ids)That loop is the whole generation process. A chatbot reply of 300 tokens is 300 runs of the model, each one predicting a single token.
03Where does the model's knowledge come from?
From training. During pretraining the model reads trillions of tokens and, at every position, predicts the next one. Each wrong prediction nudges the weights slightly so the right token becomes more likely next time. Repeated over an enormous corpus, this forces the model to absorb grammar, facts, styles, reasoning patterns and code, because all of them help predict text.
Then it is fine-tuned to follow instructions and behave helpfully. Part 4 walks through each stage.
The important consequence: knowledge lives in the weights, compressed and approximate. The model does not look facts up. It produces text that is statistically likely given what it learned, which is usually right and sometimes confidently wrong. That is the root of hallucination, the subject of part 6.
04What does the model not have?
Three things beginners often assume it has:
- Memory between requests. Each request starts fresh. A chat app feels continuous because it re-sends the conversation history every time.
- Access to live data. Unless the application fetches information and puts it in the prompt, the model only knows what was in its training data, up to a cutoff date.
- The ability to act. A model only produces text. When it “searches the web” or “books a meeting”, an application around it is reading the model's output and calling real tools. That is what an agent is, starting in part 8.
An LLM is a next-token predictor. Chat, code, search and agents are all applications built around that one capability.
05What is a context window?
The context window is the maximum number of tokens the model can take in one request: the prompt plus everything generated so far. Modern models handle tens of thousands to hundreds of thousands of tokens. Everything the model can use to answer must fit inside it, which is why applications spend so much effort choosing what to put there.
06Where to go next
Part 2 zooms into the first two steps of the pipeline: how text is split into tokens, and how tokens become vectors that carry meaning.