AGENSPHERE/ JOURNAL
← JOURNAL
FOUNDATIONS · PART 10 OF 10
J-024GUIDE4 MIN READ

Agent memory, planning and stopping: what turns a loop into a system

A basic agent remembers only its message list, plans one step at a time and stops at an arbitrary limit. Real agents manage their memory, plan when the task is long, and stop on purpose. Here is how each part works, and where to go from here.

IN SHORT
  • Agent memory comes in two kinds: working memory (the context of the current run, kept within a budget) and long-term memory (facts and summaries stored between runs and retrieved when relevant).
  • Planning helps on long tasks: the agent writes a plan, executes steps, and revises the plan when results surprise it. Short tasks rarely need an explicit plan.
  • Multi-agent designs split a task between specialised agents. They help when subtasks are truly separable and cost more in coordination, tokens and debugging.
  • Agents should stop on purpose: when the goal is met, when they are stuck, when a budget is spent, or when a person must decide. Each stop should leave a clear, logged outcome.

This is part 10, the last part of Foundations. In part 9 we built an agent in about twenty lines. It worked, but it remembered only its message list, thought one step at a time and stopped at an arbitrary limit. This part covers the three things that separate a demo loop from an agent you can rely on.

01How does AI agent memory work?

Recall from part 1 that the model itself remembers nothing between calls. All memory is something the system around the model builds. There are two kinds.

Working memory is what the agent knows during the current run: the task, the steps taken, the tool results. In the basic loop it is the full message history, re-sent on every step. That works for short runs. On long runs the history grows until it is expensive, slow and too long for the model to use well.

Ways to keep working memory in budget:

  • Trim tool results to what matters before adding them to the history.
  • Summarise older steps into a compact running record, keeping recent steps in full.
  • Keep a scratchpad of key facts (“order is paid, refund limit is $200”) separate from the raw transcript.

Long-term memory is what persists between runs: a user's preferences, the outcome of a previous ticket, facts learned about an account. It is stored outside the model, usually as structured records or embedded summaries, and retrieved when relevant, the same way RAG retrieves documents (see part 7).

The engineering choices are what to write, when to update it, and when to forget. We go deeper in a context window is not memory.

02How does an agent plan?

The basic loop plans implicitly: at each step, the model picks the next action. That is fine for tasks of a few steps. For longer tasks, such as researching a question across many sources or migrating a set of records, an explicit plan helps:

  1. Plan: the model writes a short list of steps toward the goal.
  2. Execute: it works through the steps, calling tools.
  3. Re-plan: when a result is surprising (a record is missing, an API fails), it revises the remaining steps instead of pushing on.
agent/plan_execute.py
def run_with_plan(goal: str, tools, max_steps=25):
    plan = llm.structured(PlanSchema, f"Goal: {goal}\nWrite 3 to 7 concrete steps.").steps
    done, notes = [], []
    for _ in range(max_steps):
        if not plan:
            return llm.chat(final_prompt(goal, done, notes)).text
        step = plan.pop(0)
        result = run_agent(step, tools, max_steps=5)          # the loop from part 9, scoped to one step
        done.append((step, result))
        check = llm.structured(ReviewSchema, review_prompt(goal, done, plan))
        if check.revise:
            plan = check.new_plan                             # adapt instead of pushing on
        notes += check.facts                                  # scratchpad: compact working memory
    return "Stopped: step budget spent. Partial results attached."

Two cautions: plans written before any information is gathered are often wrong, so re-planning matters more than the first plan. And planning adds model calls, so use it where tasks are long enough to need it.

03Should you use multiple agents?

Multi-agent systems split a task between several agents, each with its own instructions and tools: a researcher, a writer and a reviewer, or one agent per system it needs to work in. An orchestrator hands out subtasks and combines results.

They help when subtasks are genuinely separable and benefit from narrow instructions and tool sets. They cost more: extra model calls, information lost in hand-offs between agents, and much harder debugging when something goes wrong. A single agent with well-designed tools is often enough, and is always easier to understand.

04How should an agent stop?

The basic loop stops when the model stops calling tools, or at a step limit. Production agents need deliberate stopping rules, each with a clear outcome:

Stop reasonOutcome
Goal metFinal answer and a record of actions taken
Stuck (repeating the same call, no progress)Stop early, explain what was tried, hand over
Budget spent (steps, tokens, time or money)Stop with partial results, never silently
Needs a decisionPause for a person, as a durable wait, not a blocking call
Unsafe or out of scopeRefuse and log

Every stop should be logged with its reason. “The agent stopped” is not an outcome; “stopped after 12 steps because the payment API returned 503 three times, refund not issued” is.

A demo agent runs until it stops. A production agent stops for a reason, and tells you which one.

05What changes on the way to production?

You now know how an LLM works, how it is trained, why it hallucinates, how RAG grounds it, and how tools, loops, memory and planning turn it into an agent. Production adds engineering around all of it, which is what the rest of this journal is about:

That is the gap between an agent that works in a demo and a system that holds up in production. Closing it is the work we do.

Previous: How AI agents work. Back to the start: How a large language model works.

Building something like this?DESCRIBE A SYSTEM →