AGENSPHERE/ JOURNAL
← JOURNAL
J-005DECISION3 MIN READ

Trace and cost every model call from day one

If a request gets slow or expensive and you cannot say which step did it, you cannot fix it. LLM observability starts with one span per model and tool call, with tokens and cost attached: cheap to add early, painful to add late.

IN SHORT
  • Tracing an AI system means recording one span per model call and tool call, linked into a single trace per user request.
  • Each model span should carry the model, prompt version, input and output tokens, cost, latency and a cache-hit flag, so cost can be grouped by feature, customer or step.
  • OpenTelemetry works for this: use its existing trace model, add model-specific attributes, and send spans to whatever backend you already run.
  • Adding it on day one costs a decorator. Adding it after launch means guessing which step caused a cost spike that already happened.

The first version of an AI feature usually has one log line: the final answer. Then the bill arrives, or a customer reports a 40 second response, and nobody can say which of the nine model calls inside that request was responsible.

01The decision

DECISION RECORD · DR-03LAYER L0
CONTEXTA single user request fans out into several model calls, retrieval queries and tool calls. Costs and latency were only visible as a monthly provider invoice and an average response time.
OPTIONSLog the final request and response only · Rely on the model provider's usage dashboard · Emit a trace with one span per model and tool call, carrying tokens, cost and latency
CHOSENEvery model call and tool call is a span in one trace per request. Spans carry tokens, cost and the prompt version, and roll up by feature and customer.
REJECTEDRequest-level logs cannot attribute cost to a step. Provider dashboards show spend by API key, not by feature or customer, and arrive too late to debug a single slow request.

02What goes on a model call span?

Enough to answer “why was this slow?” and “why did this cost that much?” without opening another tool.

AttributeWhy it matters
model and providerPrice and latency differ by an order of magnitude between models
prompt.versionTies a cost or quality change to a specific prompt commit
tokens.input, tokens.outputCost is computed from these, and long inputs are the usual culprit
tokens.cachedPrompt caching changes the price; you want to see whether it hit
cost.usdComputed at call time from a price table you keep in the repo
latency.ms, ttft.msTotal time and time to first token are different problems
feature, customer, stepThe dimensions you will group by when the bill surprises you

Tool calls and retrieval queries get spans too, with their own latency and result counts. The whole request is one trace, so the slow step is visible as the longest bar.

03How do you add LLM observability without a platform?

You probably already run tracing for your services. OpenTelemetry has a trace model that fits: a span per call, attributes for the details, a trace ID that follows the request. Wrap the model client once.

observability/llm.py
from opentelemetry import trace
tracer = trace.get_tracer("llm")

def call_model(step: str, model: str, messages, prompt_version: str, **params):
    with tracer.start_as_current_span(f"llm {step}") as span:
        span.set_attributes({"llm.model": model, "llm.step": step,
                             "prompt.version": prompt_version})
        resp = client.chat(model=model, messages=messages, **params)
        u = resp.usage
        span.set_attributes({
            "tokens.input": u.input_tokens,
            "tokens.output": u.output_tokens,
            "tokens.cached": getattr(u, "cached_tokens", 0),
            "cost.usd": price(model, u),        # price table lives in the repo
        })
        return resp

Set feature and customer once at the start of the request as baggage or span attributes, so every child span carries them.

04What do you do with the data?

Three views pay for the effort almost immediately:

  1. Cost per request, by feature. Not the monthly total. The cost of one “summarise this thread” is the number product decisions need.
  2. The slowest step, by percentile. Averages hide the long tail. Look at p95 per step, and the culprit is usually one retrieval call or one oversized prompt.
  3. Tokens per step over time. A context that grows release after release shows up here weeks before it shows up on the invoice.

You cannot optimise a cost you can only see as a monthly total.

05Is this worth it for a prototype?

Yes, because it is cheapest at the start. A wrapper around the model client is an afternoon. Retrofitting it later means touching every call site, and the data you needed to explain last month's spike is already gone.

It is also why our durable execution runtime, KEEL, records tokens, cost and latency on every step in its log: the economics of a run should come out of the same record as its behaviour. This sits in layer L0 of our stack: infrastructure, including observability and cost.

Building something like this?DESCRIBE A SYSTEM →