- Tracing an AI system means recording one span per model call and tool call, linked into a single trace per user request.
- Each model span should carry the model, prompt version, input and output tokens, cost, latency and a cache-hit flag, so cost can be grouped by feature, customer or step.
- OpenTelemetry works for this: use its existing trace model, add model-specific attributes, and send spans to whatever backend you already run.
- Adding it on day one costs a decorator. Adding it after launch means guessing which step caused a cost spike that already happened.
The first version of an AI feature usually has one log line: the final answer. Then the bill arrives, or a customer reports a 40 second response, and nobody can say which of the nine model calls inside that request was responsible.
01The decision
02What goes on a model call span?
Enough to answer “why was this slow?” and “why did this cost that much?” without opening another tool.
| Attribute | Why it matters |
|---|---|
model and provider | Price and latency differ by an order of magnitude between models |
prompt.version | Ties a cost or quality change to a specific prompt commit |
tokens.input, tokens.output | Cost is computed from these, and long inputs are the usual culprit |
tokens.cached | Prompt caching changes the price; you want to see whether it hit |
cost.usd | Computed at call time from a price table you keep in the repo |
latency.ms, ttft.ms | Total time and time to first token are different problems |
feature, customer, step | The dimensions you will group by when the bill surprises you |
Tool calls and retrieval queries get spans too, with their own latency and result counts. The whole request is one trace, so the slow step is visible as the longest bar.
03How do you add LLM observability without a platform?
You probably already run tracing for your services. OpenTelemetry has a trace model that fits: a span per call, attributes for the details, a trace ID that follows the request. Wrap the model client once.
from opentelemetry import trace
tracer = trace.get_tracer("llm")
def call_model(step: str, model: str, messages, prompt_version: str, **params):
with tracer.start_as_current_span(f"llm {step}") as span:
span.set_attributes({"llm.model": model, "llm.step": step,
"prompt.version": prompt_version})
resp = client.chat(model=model, messages=messages, **params)
u = resp.usage
span.set_attributes({
"tokens.input": u.input_tokens,
"tokens.output": u.output_tokens,
"tokens.cached": getattr(u, "cached_tokens", 0),
"cost.usd": price(model, u), # price table lives in the repo
})
return respSet feature and customer once at the start of the request as baggage or span attributes, so every child span carries them.
04What do you do with the data?
Three views pay for the effort almost immediately:
- Cost per request, by feature. Not the monthly total. The cost of one “summarise this thread” is the number product decisions need.
- The slowest step, by percentile. Averages hide the long tail. Look at p95 per step, and the culprit is usually one retrieval call or one oversized prompt.
- Tokens per step over time. A context that grows release after release shows up here weeks before it shows up on the invoice.
You cannot optimise a cost you can only see as a monthly total.
05Is this worth it for a prototype?
Yes, because it is cheapest at the start. A wrapper around the model client is an afternoon. Retrofitting it later means touching every call site, and the data you needed to explain last month's spike is already gone.
It is also why our durable execution runtime, KEEL, records tokens, cost and latency on every step in its log: the economics of a run should come out of the same record as its behaviour. This sits in layer L0 of our stack: infrastructure, including observability and cost.