AGENSPHERE/ JOURNAL
← JOURNAL
J-013NOTE4 MIN READ

Cut LLM cost in the right order: cache, route, batch, then trim

Most attempts to reduce LLM cost start with prompt golf and end with a quality regression. Measure first, then take the structural wins: prompt caching, model routing and batch processing, before rewriting a single prompt.

IN SHORT
  • LLM cost is input tokens plus output tokens times the price of the model, summed over every call in a request. You cannot reduce it sensibly until you can see it per step.
  • Prompt caching discounts repeated input prefixes. Put stable content (system prompt, tool definitions, reference documents) first and variable content last so the prefix can be reused.
  • Model routing sends easy, high-volume steps to a small model and keeps the large model for hard ones, checked against an eval so quality does not slip.
  • Batch interfaces process work that does not need an instant answer at a large discount. Trimming prompts comes last, because it saves the least and risks the most.

When an LLM bill grows, the first instinct to reduce LLM cost is to shorten prompts. It is the least effective lever and the riskiest one, because every word removed is a possible behaviour change. There is a better order.

01Step zero: can you see the cost per step?

LLM cost for one request is the sum over every model call: input tokens times the input price, plus output tokens times the output price, for whichever model each call used. A request that fans out into a planner, three tool-using turns and a summariser has five line items, not one.

If you only see a monthly total, you are guessing. Trace every call with its model, token counts and computed cost, grouped by feature and step. We cover that in trace and cost every model call from day one. The rest of this note assumes you can answer “which step costs the most?”

021. Cache: stop paying full price for the same prefix

Most providers now discount input tokens that repeat a prefix they have recently seen, either automatically or with explicit cache markers. Agents are full of repeated prefixes: the same system prompt, the same tool definitions, the same reference documents on every turn.

To benefit, order the prompt from most stable to least stable:

agent/prompt.py
def build_messages(task, history, docs):
    return [
        system(SYSTEM_PROMPT),            # identical on every call: cached
        system(render_tools(TOOLS)),      # changes only on deploy: cached
        system(render_docs(docs)),        # stable within a session: often cached
        *history,                         # grows each turn: partly cached
        user(task.text),                  # new every time: never cached
    ]

The classic mistake is putting something variable at the top: a timestamp, a request ID, the user's name in the first line of the system prompt. That one token breaks the prefix and every call pays full price. Check your traces for the cached-token count; if it is near zero on a multi-turn agent, the prefix is unstable.

032. Route: use the smallest model that passes the eval

Not every step needs the most capable model. Classifying a ticket, extracting fields, choosing which tool to call next and summarising a short thread are often handled well by a much smaller and cheaper model. Hard reasoning, ambiguous requests and final customer-facing replies may need the large one.

Route per step, not per product, and let an eval decide:

StepVolumeCandidateKeep it if
Ticket classificationHighSmall modelAccuracy on the eval set stays within a point of the large model
Field extractionHighSmall model with structured outputSchema and rule checks pass at the same rate
Multi-step resolutionMediumLarge model(default)
Customer replyMediumLarge model(default)

Run the eval on both models for each step before switching. Routing without an eval is how cost drops and complaints rise. See ship a prompt change with an eval.

043. Batch: not everything needs an answer in two seconds

A lot of LLM work is not interactive: nightly summaries, backfills, re-classifying old tickets, generating embeddings, running evals. Major providers offer batch interfaces that accept large sets of requests, return results within hours, and charge substantially less, commonly around half the normal price.

Sort your workload into “a user is waiting” and “nobody is waiting”. Everything in the second group is a batch candidate.

054. Trim: shorten what is left, carefully

Only now look at the prompts themselves, starting with the most expensive step from your traces:

  • Retrieval budget. Sending eight retrieved chunks when three answer the question is often the biggest waste. Re-rank and cut.
  • History. Replace old turns with a summary instead of re-sending the full transcript. See a context window is not memory.
  • Output length. Output tokens usually cost several times more than input tokens. Ask for the shape you need, and set a sensible maximum.
  • Tool definitions. Long descriptions for tools a step never uses are paid on every call. Give each step only its tools.

Every trim goes through the eval like any other prompt change.

Structure saves money without touching behaviour. Trimming saves money by touching it. Do the first before the second.

06What to report

A useful cost report is short: cost per request by feature, cost per step for the top features, cache hit rate, the share of calls on each model, and the share of work running in batch. Watch those over time and cost stops being a surprise.

This sits across L0 infrastructure, where tracing and cost live, and L2 reasoning, where model choice and prompts live.

Building something like this?DESCRIBE A SYSTEM →