- Tool calling lets a model ask for an action. You describe tools (name, description, input schema); the model replies with a structured request naming a tool and its arguments.
- The model never executes anything. Your application runs the tool, sends the result back as a message, and the model continues with that new information.
- Good tool design matters as much as prompting: clear names and descriptions, tight schemas, and results that are short and informative, including errors.
- Because your code runs every call, your code is where permissions, validation and approvals belong.
This is part 8 of Foundations. So far the model has only read and written text. Part 7 fed it documents. Now we let it ask for things: look up an order, query a database, send a message. This is tool calling, and it is the bridge from a chatbot to an agent.
01What is tool calling?
Tool calling, also called function calling, is a feature of modern LLM APIs. Along with your messages, you send a list of tools the model may use. Each tool has:
- a name, like
lookup_order - a description of what it does and when to use it
- an input schema, usually JSON Schema, describing the arguments
When the model decides a tool would help, instead of (or as well as) writing a normal reply it returns a tool call: the tool's name and arguments as structured data.
02Who actually runs the tool?
You do. This is the most important point in the whole topic. The model is a text predictor (see part 1); it has no ability to execute code, open a network connection or touch a database. Tool calling is a protocol where the model asks and your application does.
One full round trip looks like this:
- You send the conversation plus tool definitions.
- The model returns a tool call:
lookup_order({"order_id": "ORD-1002"}). - Your code validates the arguments, checks the user may do this, and runs the real function.
- You send the result back as a tool-result message.
- The model reads the result and either calls another tool or writes its final answer.
TOOLS = [{
"name": "lookup_order",
"description": "Look up an order by id. Use before answering any question about an order's status, total or payment.",
"input_schema": {
"type": "object",
"properties": {"order_id": {"type": "string", "description": "Order id, like ORD-1002"}},
"required": ["order_id"],
},
}]
messages = [{"role": "user", "content": "Where is my order ORD-1002?"}]
reply = llm.chat(messages=messages, tools=TOOLS)
for call in reply.tool_calls: # the model asked; now we act
if call.name == "lookup_order":
result = orders.get(call.args["order_id"], scope=current_user)
messages += [reply.as_message(), tool_result(call.id, result)]
final = llm.chat(messages=messages, tools=TOOLS) # the model answers using the resultExact field names differ between providers, but every major API follows this shape.
03How does the model know when to call a tool?
It was trained to. Instruction tuning and preference tuning (see part 4) include many examples of conversations with tools, so the model learns to emit a well-formed call when the description matches the need. The tool definitions are placed into the model's context, so they are also tokens it reads and pays for on every request.
That is why the description is part of your prompt. A vague description (“gets data”) leads to wrong or missed calls; a precise one (“use before answering any question about an order's status”) leads to reliable behaviour.
04What makes a good tool?
- One clear job per tool.
lookup_orderandissue_refund, notmanage_orders(action=...). - Names and descriptions written for the model. Say what it does, when to use it, and what it returns.
- Tight schemas. Enums for fixed choices, formats and ranges for values, required fields marked. Every constraint is a mistake the model cannot make.
- Short, informative results. Return what the model needs, not a 5,000-token raw API response.
- Errors the model can act on. “Order ORD-1002 not found; check the id with the customer” lets the model recover; a stack trace does not.
05Where does safety belong?
In your code, around every call, because that is the only place that is guaranteed to run. The model's choice of tool and arguments is a request from a component that can be wrong or manipulated (see prompt injection).
So the application should:
- run each tool with the acting user's permissions, never a shared admin account,
- validate arguments against the schema and business rules before running anything,
- know each tool's effect class: whether running it twice is harmless or a second refund (see effect classes),
- require approval for irreversible actions.
The model proposes a function call. Your system decides whether it happens.
06What about MCP?
The Model Context Protocol standardises how tools are described and connected, so one integration works across many AI applications. It changes how tools are packaged, not who runs them or who is responsible for permissions. See MCP turns tools into a protocol.
07Where to go next
One tool call answers one question. Put tool calling inside a loop, where the model keeps calling tools and reading results until the task is done, and you have an agent. Part 9 walks through that loop.
Previous: How RAG works. Next: How AI agents work: the loop.