- A context window is the text a model reads on one request. It is rebuilt from scratch every time and holds nothing between calls.
- Memory is state that persists between requests and is chosen on purpose: what to keep, for how long, and when to bring it back.
- Stuffing history into the window grows cost and latency with every turn, and models use the middle of a long context less reliably than the start and end.
- Engineer memory explicitly: structured facts, summaries with sources, and retrieval with a budget per request, plus a rule for forgetting.
Context windows have grown from a few thousand tokens to hundreds of thousands. It is tempting to read that as “the model can now remember the whole conversation”. It cannot. It can read the whole conversation again, on every request, and you pay for every token each time.
01What is the difference?
A context window is the input to one model call. Your application assembles it, the model reads it, and it is gone when the call returns. The model keeps nothing.
Memory is state your system stores between calls and decides to show the model again. It has a lifecycle: something is written, kept for a reason, retrieved when relevant, and eventually removed.
The model is stateless either way. Memory is something you build around it.
02Why not just send everything?
It works in a demo and then stops working.
Cost grows with every turn. If each request re-sends the full history, turn 50 pays for 49 turns of text again. The total cost of a conversation grows roughly with the square of its length.
Latency grows with it. More input tokens mean a longer time to first token, on every turn.
Long contexts are used unevenly. Models tend to use information at the start and end of a long input more reliably than information buried in the middle. A fact said once, forty turns ago, is present but easy to miss.
Stale and wrong facts stay forever. If the user corrected an address in turn 12, the old address from turn 3 is still in the window, and the model has to work out which one wins.
03What does engineered LLM memory look like?
Decide what kinds of memory the system needs, and give each one a store and a rule.
| Kind | What it holds | How it comes back |
|---|---|---|
| Working | The current task: last few turns, open tool results | Always in the window, trimmed by a budget |
| Facts | Stable statements: “prefers email”, “account is on the EU region” | Structured lookup by user or entity |
| Episodic | Summaries of past sessions, with links to the source turns | Retrieved when relevant to the current request |
| Reference | Documents and knowledge bases | Retrieval, filtered by permission |
Then assemble each request on purpose, inside a token budget:
def build_context(user, task, budget=12_000):
parts = [
system_prompt(), # fixed, cached
facts.for_user(user).render(), # small, structured
episodes.search(task.query, user=user, k=3), # summaries with sources
docs.search(task.query, scope=user.scope, k=6), # permission-filtered
task.recent_turns(max_tokens=4_000), # working memory
]
return fit_to_budget(parts, budget, drop_order=["docs", "episodes"])04Writing memory is a decision too
Most systems are careful about reading memory and careless about writing it. Rules worth having:
- Write facts, not transcripts. Extract “the user's order ORD-1002 was refunded” as a fact with a source, rather than storing the whole exchange.
- Overwrite, do not append. A corrected fact replaces the old one, with a timestamp.
- Keep the source. Every fact and summary links back to the turn or document it came from, so it can be checked and deleted.
- Forget on purpose. Expiry rules and user deletion requests apply to memory like any other personal data.
More context is more reading. Memory is choosing what is worth reading again.
05Where this is heading
Everything above is how to work well with stateless models today: memory as an engineering layer around a model that starts from zero on every request. It sits in layer L3 of our stack, intelligence: memory, context and decision policy.
There is a deeper question underneath: what if the model itself kept state and learned from experience, instead of re-reading its history? That is the research direction of June Labs, Agensphere's research arm, at junelabs.agensphere.com. Until that exists, the practical answer is the one in this note: treat the context window as a budget, and build memory deliberately.