- Build vs buy AI goes wrong when it is treated as one decision for a whole system. Decide per layer: infrastructure, data and knowledge, reasoning, intelligence, actions and workflow.
- Buy, or rent, what is commodity and improves without you: foundation models, cloud infrastructure, vector databases, and AI tools for work that does not differentiate you.
- Build and own what encodes how your business works: the context layer over your data, evaluation sets, routing and policy, integrations with your systems of record, and the audit trail.
- Building has a cost buying hides: someone has to own, run and improve it. Only build what you are prepared to operate.
Build vs buy AI sounds like a single question with a single answer. In practice, almost every serious AI system is part bought and part built, and the expensive mistakes come from drawing the line in the wrong place: building what you should have rented, or renting what you needed to own.
01Why is build vs buy AI the wrong single question?
An AI system is a stack, and each layer has a different answer. We describe systems in six layers, from L0 infrastructure to L5 product and workflow. Asked layer by layer, the build vs buy decision gets much easier:
| Layer | Usually buy or rent | Usually build and own | Why |
|---|---|---|---|
| L0 Infrastructure | Cloud, GPUs, managed databases | Tracing, cost per task, routing config | Compute is a commodity; knowing what it is spent on is not |
| L1 Data / Knowledge | Vector database, search engine | Connectors, chunking, permissions, the index | Your data and who may see it is unique to you |
| L2 Reasoning | Foundation models | Prompts, evaluation sets, guardrails | Models improve without you; your definition of "good" does not |
| L3 Intelligence | Agent frameworks, sparingly | Memory, decision policy | What the system may do on its own is a business decision |
| L4 Actions | Standard connectors, MCP servers | Tool permissions, approvals, idempotency | Actions touch your systems of record |
| L5 Product / Workflow | Off-the-shelf AI for generic work | The workflow your team actually runs | Fit to the work is where adoption is won or lost |
02What should you buy?
Buy what is commodity, improves without your help and would be expensive to keep up with yourself:
- Foundation models. Almost nobody should train one. Rent them through APIs, keep them swappable (see avoid LLM vendor lock-in) and route each task to the cheapest one that passes its eval.
- Infrastructure. Cloud compute, managed databases and vector search are mature products.
- AI tools for non-differentiating work. Meeting notes, coding assistants, writing help. If a tool does a generic job well and your data rules allow it, buy seats.
03What should you build and own?
Own the parts that encode how your business works, or that you would lose if a vendor changed its terms:
- The context layer: connectors to your systems, chunking that respects your documents, and permissions enforced in retrieval (see out of context enterprise AI).
- Evaluation sets: your definition of a good answer, as tests. They are what let you switch models safely.
- Policy, routing and the audit trail: which data may go where, which model handles which task, and a record of what the system saw and did (see AI audit trail).
- Workflow integrations: the AI inside the tools your people already use, not in another tab.
If a vendor holds your prompts, your evals and your logs, you have not bought a tool. You have rented your own intelligence back (see unowned intelligence).
04Which questions decide each case?
# Answer per component, not per system.
component: claims-triage-assistant
questions:
differentiates_us: yes # does doing this better win customers or margin?
uses_sensitive_data: yes # can the data leave our boundary under a vendor's terms?
fits_existing_workflow: no # does an off-the-shelf tool work inside our claims system?
high_volume: yes # will per-seat or per-call pricing dominate at scale?
switching_cost_if_vendor_changes: high
we_can_operate_it: yes # is there a named owner and on-call after launch?
verdict: build the context, policy and workflow layers; rent the models and the vector storeTwo or more "yes" answers in the first four questions usually point to building that component. A "no" on the last one means do not build it yet, whatever the rest say.
05What does building cost that buying hides?
Ownership. A built component needs a named owner, monitoring, on-call, a budget for model and infrastructure costs, and someone who keeps its eval set current as the business changes. Bought software includes that work in the price. Systems that are built and then left without an owner are how pilots stall before production.
So the honest rule is: build what differentiates you and what you must control, and only as much of it as you are prepared to run.
Rent the intelligence that improves without you. Own the intelligence that only you can define.
06Where to go next
For the full architecture of the owned parts, read the owned intelligence layer. For who should build them, read AI engineering partner vs in-house team. The terms used here are defined in the glossary.