AGENSPHERE/ JOURNAL
← JOURNAL
FOUNDATIONS · PART 2 OF 10
J-016GUIDE4 MIN READ

Tokens and embeddings: how text becomes numbers

A model never sees letters. It sees token IDs, and then vectors. Knowing what a token is in an LLM explains pricing, context limits and odd failures with spelling and numbers; embeddings explain how a model represents meaning.

IN SHORT
  • A token is the unit a model reads and writes: often a whole common word, a piece of a rarer word, or a punctuation mark. In English, one token averages roughly four characters, or about three quarters of a word.
  • Tokenisers such as byte pair encoding (BPE) build their vocabulary by repeatedly merging the most frequent character pairs, so common words become one token and rare words split into several.
  • Each token ID is turned into an embedding: a vector of numbers learned during training, where related meanings end up close together.
  • Tokens explain pricing and context limits (both are counted in tokens) and some quirks, like models miscounting letters, because the model sees chunks, not characters.

This is part 2 of Foundations. In part 1 we followed a request through the model. Here we slow down on the first two steps: turning text into tokens, and turning tokens into vectors.

01What is a token in an LLM?

A neural network works on numbers, so text has to be converted before the model can use it. The conversion happens in chunks called tokens.

A token is usually one of:

  • a whole common word, often including the space before it: “ the”, “ model”
  • a piece of a longer or rarer word: “token” + “isation”
  • punctuation or a symbol: “.”, “{”
  • part of a number: “2026” might be one token or two

As a rule of thumb for English, one token is about four characters, or about three quarters of a word. Other languages, code and numbers often take more tokens for the same amount of meaning.

02How does a tokeniser decide where to split?

Most modern models use a subword method such as byte pair encoding (BPE) or a close relative. The idea is simple:

  1. Start with a vocabulary of single bytes or characters.
  2. Look through a large training corpus and find the pair of adjacent symbols that appears most often.
  3. Merge that pair into a new symbol and add it to the vocabulary.
  4. Repeat until the vocabulary reaches a target size, often 50,000 to a few hundred thousand.

Frequent sequences like “ing” and “ the” get merged early and become single tokens. Rare words never get merged fully, so they are spelled out from smaller pieces. This gives the model a fixed vocabulary that can still represent any text.

tokens_demo.py
import tiktoken                                   # OpenAI's open-source tokenizer library

enc = tiktoken.get_encoding("cl100k_base")
for text in ["The model", "tokenisation", "ORD-1002", "strawberry"]:
    ids = enc.encode(text)
    print(f"{text!r:16} -> {len(ids)} tokens: {[enc.decode([i]) for i in ids]}")

Running something like this is the quickest way to build intuition. Common words come out as one token; product codes and unusual words break into several.

03Why do tokens matter in practice?

Cost and limits are counted in tokens. API pricing is per input and output token, and the context window is a token limit. A document that is 10,000 words long is roughly 13,000 tokens in English.

Models see chunks, not letters. When a model is asked how many “r”s are in “strawberry”, it is reasoning about tokens like “str” and “awberry”, not individual characters. Spelling, counting letters and some arithmetic are harder than they look for exactly this reason.

Non-English text costs more. Tokenisers trained mostly on English need more tokens for many other languages, which means higher cost and less room in the context window for the same content.

04What is an embedding?

A token ID like 4089 is just a label. The number itself means nothing. The model's first real step is to look that ID up in an embedding table: a large matrix with one row per vocabulary entry. Each row is a vector of a few thousand numbers.

These vectors are learned during training. Nobody assigns them by hand. Because the model is trained to predict text, tokens used in similar contexts drift toward similar vectors. “cat” and “dog” end up closer to each other than to “carburettor”.

You can think of an embedding as coordinates in a space with thousands of dimensions, where direction and distance carry meaning. Closeness is usually measured with cosine similarity, the angle between two vectors.

similarity.py
import numpy as np

def cosine(a: np.ndarray, b: np.ndarray) -> float:
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))

# With real embeddings, cosine(cat, dog) is noticeably higher than cosine(cat, invoice).

05How does the model know word order?

Embeddings on their own say nothing about position: “dog bites man” and “man bites dog” contain the same tokens. Transformers add positional information so the model can tell order. Older models added a fixed position vector to each embedding; many current models use rotary position embeddings (RoPE), which encode relative position inside the attention calculation itself.

06Token embeddings vs embedding models

There are two related things called embeddings, and it helps to keep them apart:

Token embeddingsText embeddings (embedding models)
WhatOne vector per token, inside an LLMOne vector for a whole sentence or passage
Used forThe first step of every forward passSearch, clustering, retrieval (RAG)
Where you meet themRarely, they are internalVector databases, semantic search

Embedding models are trained specifically so that passages with similar meaning get similar vectors. They power the retrieval step in RAG, which we walk through in part 7.

The model never reads your words. It reads token IDs, turns them into vectors, and does all its work on those vectors.

07Where to go next

Once each token is a vector, the transformer layers take over. Part 3 explains attention, the mechanism that lets each token use context from the others.

Previous: How a large language model works. Next: Attention and the transformer, explained.

Building something like this?DESCRIBE A SYSTEM →