- An AI agent is an LLM that decides which tools to call, sees the results, and repeats until it can finish the task. The core is a loop around tool calling.
- Every agent has the same parts: a model, instructions, tools, a running state (the message history), and a stop condition.
- The model chooses the path at runtime. That makes agents flexible on open-ended tasks and harder to predict, test and bound on cost.
- When the steps are known in advance, a fixed workflow that calls the model at defined points is usually cheaper, faster and more reliable than an agent.
This is part 9 of Foundations. Part 8 showed one round of tool calling: the model asks, your code runs the tool, the model reads the result. An agent is what you get when you let that happen as many times as the task needs.
01What is an AI agent?
The word is used loosely, so here is a precise definition: an AI agent is a system in which a language model decides, step by step, which actions to take to reach a goal, using tools, and keeps going until it decides the goal is met.
The key word is decides. In ordinary software, a developer writes the sequence of steps. In an agent, the model chooses the next step at runtime, based on what it has seen so far.
02How do AI agents work? The five parts
Every agent, however it is branded, has five parts:
- A model that reads the situation and chooses the next action.
- Instructions: the goal, the rules, and how to behave (the system prompt).
- Tools it can call to read information or change things (see part 8).
- State: what has happened so far. In the simplest agents this is just the growing list of messages and tool results.
- A stop condition: the model gives a final answer, or a limit on steps, time or cost is reached.
03What does the loop look like?
Here is a complete agent. It is short on purpose: the loop really is this simple.
def run_agent(task: str, tools: dict, max_steps: int = 10) -> str:
messages = [system(INSTRUCTIONS), user(task)]
for step in range(max_steps):
reply = llm.chat(messages=messages, tools=[t.schema for t in tools.values()])
messages.append(reply.as_message())
if not reply.tool_calls: # no more actions: the model is done
return reply.text
for call in reply.tool_calls: # the model asked for actions; we run them
tool = tools.get(call.name)
try:
result = tool.run(**call.args) if tool else {"error": f"unknown tool {call.name}"}
except ToolError as e:
result = {"error": str(e)} # errors go back to the model, not up the stack
messages.append(tool_result(call.id, result))
return "Stopped: step limit reached. Escalating to a person."Walk through a support request with it. The user asks for a refund on a damaged order.
- Step 1. The model reads the task and calls
lookup_order("ORD-1002"). - Step 2. It sees the order is paid, $120, and calls
issue_refund(order_id="ORD-1002", amount_cents=12000). - Step 3. It sees the refund succeeded and writes a short reply to the customer. No tool calls, so the loop ends.
Nobody wrote “look up, then refund, then reply”. The model chose that path from the instructions, the tool descriptions and what each result told it.
04Why does reasoning between steps help?
Models make better choices when they think before acting. A well-known pattern called ReAct (reason and act) has the model alternate between a short piece of reasoning (“the order is paid and the amount is within the total, so I can refund”) and an action, then observe the result. Modern models often do this internally, but the principle holds: give the model room to reason about what it just saw before choosing the next tool.
05What makes agents hard in practice?
The same freedom that makes them flexible:
- Unpredictable paths. The same task can take three steps or eleven. Testing has to cover behaviour, not one fixed sequence.
- Compounding errors. A small mistake in step 2 shapes every step after it.
- Unbounded cost and time. Each step is a model call. Without limits, a confused agent can loop expensively.
- Real side effects. Tools change the world. A retried step can mean a second refund (see effect classes).
- Long runs and crashes. A run that waits for a person or calls slow APIs may outlive the process it started in (see human approval is a durable wait).
06When should you not use an agent?
When you already know the steps. Many tasks described as “agentic” are really fixed workflows: extract fields from an invoice, check them against the purchase order, draft a reply. Hard-coding that sequence and calling the model at each defined step is cheaper, faster, easier to test and easier to explain.
| Workflow | Agent | |
|---|---|---|
| Who decides the steps | The developer, in code | The model, at runtime |
| Best for | Known, repeatable processes | Open-ended tasks with many possible paths |
| Predictability | High | Lower |
| Cost per task | Fixed and low | Variable |
A good rule: start with the simplest thing that works, often a single model call or a workflow, and move to an agent only where the task genuinely needs the model to choose its own path.
An agent is a loop. Everything that makes it production-grade is about controlling what happens inside that loop.
07Where to go next
Our simple agent remembers only its message list, has no plan, and stops at a hard step limit. Part 10, the last in the series, covers how real agents handle memory, planning and knowing when to stop.
Previous: How tool calling works. Next: Agent memory, planning and stopping.