- Enterprises often run every AI task through one top-tier model, because that is what the first prototype used or what the enterprise plan includes.
- Most production calls are routine: classify, extract, route, summarise. Smaller models handle many of these well at a fraction of the price, often an order of magnitude cheaper per token.
- A model router sends each task type to the cheapest model that meets a measured quality bar, escalates hard cases to a stronger model, and falls back when a provider fails.
- Routing must be decided by evals on your own data, not by benchmark leaderboards, and revisited when models or prices change.
This is part 1 of Enterprise AI, a series about the problems companies hit when they adopt AI at scale, and what to build instead. We start with the one that shows up on the invoice.
01What does the problem look like?
A team builds a prototype on the best model available. It works, so it ships. Then more features are built the same way, on the same model, because it is already approved, already integrated and already in the contract. A year later every AI call in the company, from tagging support tickets to drafting legal summaries, runs on the most capable and most expensive model on offer.
Nobody decided this. It accumulated.
02Why is it expensive?
Because most enterprise AI work is not hard. Look at the calls a typical deployment makes:
- classify a ticket, an email or a document into one of a few categories
- extract named fields from an invoice, a contract or a form
- decide which tool or team should handle a request
- summarise a short thread or a meeting note
- rewrite text into a fixed format
These tasks are high-volume and well-defined. Smaller, cheaper models often do them as well as the largest ones, especially with structured output and a good prompt. Price differences between model tiers are large, frequently ten times or more per token between a small model and a frontier one, and the routine tasks are exactly the ones that run millions of times.
Per-seat enterprise chat plans add a second layer: every employee gets the top model in a chat window, whether they use it for complex analysis or for fixing a sentence. That is a reasonable product for exploration. It is a poor way to run production workloads.
03What is LLM routing, and what does a router do?
A model router sits between your applications and the model providers. Every request goes through it. For each request it decides which model handles it, based on the task, and records what happened.
ROUTES = {
# task primary model escalate to max cost/call
"ticket.classify": ("small-fast", "mid-tier", 0.0005),
"invoice.extract": ("small-structured", "mid-tier", 0.002),
"reply.draft": ("mid-tier", "frontier", 0.01),
"case.investigate":("frontier", None, 0.08),
}
def route(task: str, payload, ctx) -> Result:
primary, escalate, budget = ROUTES[task]
result = call_model(primary, payload, ctx)
if escalate and needs_escalation(task, result): # low confidence, failed validation
result = call_model(escalate, payload, ctx, reason="escalated")
record(task=task, model=result.model, cost=result.cost, latency=result.latency, budget=budget)
return resultThree behaviours matter:
- Route by task, not by team. The unit is “invoice extraction”, not “the finance team's bot”.
- Escalate on evidence. Send a case to a stronger model when the cheap one is unsure or its output fails validation (see validate structured outputs), not by default.
- Fall back on failure. When a provider is down or rate-limited, the router retries on an equivalent model, so one outage does not stop the business.
04How do you decide which model gets which task?
With an eval, not a leaderboard. Public benchmarks measure general ability; your question is narrower: does this model do our invoice extraction at our quality bar?
For each task type:
- Collect 50 to 200 real examples with known correct outputs.
- Run every candidate model on them.
- Pick the cheapest model that meets the bar, with a margin.
- Re-run when a new model ships or prices change. Routing tables go stale.
This is the same discipline as shipping every prompt change with an eval, applied to model choice.
The question is never “which model is best?” It is “which is the cheapest model that is good enough for this task?”
05What else does routing unlock?
Once every call passes through one place, you can see cost per task, per feature and per business unit. You can cache repeated prefixes, batch work that is not urgent (see cut LLM cost in the right order), and swap a provider without touching application code. Routing is usually the first piece of the owned intelligence layer we describe at the end of this series.
Next in Enterprise AI: Out of context: why your enterprise AI doesn't know your business.