AGENSPHERE/ JOURNAL
← JOURNAL
ENTERPRISE AI · PART 6 OF 9
J-030DECISION3 MIN READ

Avoid LLM vendor lock-in with a model-agnostic seam

Models are deprecated, repriced and overtaken every few months. When provider-specific calls are scattered through your code, every change is a migration: that is LLM vendor lock-in. One thin seam, with your own prompts and evals behind it, turns switching models into a configuration change.

IN SHORT
  • LLM lock-in comes from three places: provider-specific API calls spread through the codebase, prompts tuned to one model's quirks, and features (assistants, files, workflows) that only exist on one platform.
  • Models change constantly: new releases, price changes and deprecations of older versions. A system that cannot switch models pays for that inflexibility every time.
  • A model-agnostic seam is one interface in your code for all model calls, with providers behind it as adapters, chosen per task by configuration.
  • The seam only works with evals: you can switch a model safely when you can measure, on your own data, that the replacement is as good.

This is part 6 of Enterprise AI. Part 3 argued that companies should rent models and own the intelligence around them. This part is the engineering that makes renting safe: a seam that lets you change the model without changing the system.

01Where does LLM vendor lock-in come from?

Not usually from contracts. From code and habits:

  • Calls everywhere. Provider SDK calls appear in dozens of services, each using that provider's message format, tool-calling format and error types.
  • Prompts tuned to one model. Prompts accumulate workarounds for one model's quirks and break on another.
  • Platform-only features. Hosted assistants, file stores, threads and agent builders that exist only on one provider. Convenient, and impossible to move.
  • No way to compare. Without an eval set, nobody can tell whether a different model would be as good, so nobody dares switch.

02Why does it matter now?

Because the model market moves faster than almost any dependency you have. New models arrive every few months, often better and cheaper. Prices change. Older model versions are deprecated on a published schedule, which forces a migration whether you planned one or not. Each time, a locked-in system pays the full cost of change; a flexible one changes a line of configuration and runs its evals.

03The decision

DECISION RECORD · DR-07LAYER L0
CONTEXTModel calls used one provider's SDK directly in many services. A model deprecation notice meant changes across every service, and a cheaper model for classification could not be adopted without a rewrite.
OPTIONSStay on one provider and absorb migrations as they come · Adopt a third-party framework that wraps every provider · Own a thin internal interface for model calls, with provider adapters behind it, chosen per task in configuration
CHOSENOne internal interface: messages, tools and structured output in our own types. Adapters translate per provider. Each task names its model in configuration and is covered by an eval.
REJECTEDAbsorbing migrations repeats the cost forever. A heavy framework trades one dependency for another and hides behaviour we need to see. A thin seam we own is small, testable and replaceable.

04What does the seam look like?

Keep it small. Your applications speak your types; adapters translate.

llm/interface.py
class ModelClient(Protocol):
    def complete(self, req: CompletionRequest) -> CompletionResult: ...

@dataclass
class CompletionRequest:
    task: str                         # "invoice.extract": routing and evals key off this
    messages: list[Message]           # our message type, not a provider's
    tools: list[ToolSpec] = ()
    schema: type | None = None        # structured output, validated on our side
    max_tokens: int = 1024

ADAPTERS = {"provider_a": ProviderAAdapter(), "provider_b": ProviderBAdapter(), "local": LocalAdapter()}

def complete(req: CompletionRequest) -> CompletionResult:
    model = config.model_for(req.task)                # e.g. {"provider": "provider_b", "name": "small-fast"}
    return ADAPTERS[model.provider].complete(req, model.name)

Rules that keep the seam honest:

  • No provider imports outside adapters. A lint rule enforces it.
  • Your own types for messages, tool calls, usage and errors, so the rest of the code never sees a provider format.
  • Structured output validated on your side (see validate at the boundary), so it works the same whatever the model.
  • Per-task model choice in configuration, the same table the router uses (see part 1).

05Prompts are part of the lock-in

A prompt that only works on one model is a hidden dependency. Keep prompts in your repository, prefer clear instructions over model-specific tricks, and keep a per-model variant only where the eval shows it is needed.

06The seam only works with evals

Switching a model is safe when you can show the replacement performs as well on your tasks. That is what an eval set is for (see ship a prompt change with an eval). With evals, a deprecation notice becomes a routine job: run the candidates, compare, update configuration, ship. Without them, it is a risk nobody wants to own.

Rent the model. Own the interface, the prompts and the evals. Then the market's speed works for you instead of against you.

07Where does this fit?

The seam lives in L0 infrastructure; prompt portability and evals live in L2 reasoning. Together with routing and the gateway, it forms the backbone of the owned intelligence layer.

Previous: Shadow AI. Next: Why AI pilots never reach production.

Building something like this?DESCRIBE A SYSTEM →