AGENSPHERE/ JOURNAL
← JOURNAL
ENTERPRISE AI · PART 7 OF 9
J-031NOTE3 MIN READ

Pilot purgatory: why enterprise AI pilots never reach production

The demo impressed everyone, the pilot ran for three months, and then nothing shipped. AI pilots stall on the way to production for predictable reasons: no measure of success, no owner, no integration with real systems, and no answer for risk. Each one can be designed out from day one.

IN SHORT
  • Pilot purgatory is when an AI proof of concept works well enough to impress but never becomes a production system.
  • The usual causes: success was never defined or measured, nobody owns the system after the pilot, it never connected to real data and workflows, and risk questions were left for later.
  • A production-minded pilot starts with a measurable outcome and an eval set, runs on real data with real permissions, sits inside an actual workflow, and has a named owner and a production path.
  • Smaller and real beats broader and simulated: one workflow, fully integrated, is worth more than five demos.

This is part 7 of Enterprise AI. The earlier parts covered cost, context, ownership, compliance, sprawl and lock-in. This part covers the outcome they all feed into: the pilot that never ships.

01What is pilot purgatory?

The pattern is familiar. A demo built in two weeks impresses leadership. A pilot is approved. It runs for a quarter with a small group of users, gets positive feedback, and then stalls. There is always a next step: more testing, a security review, waiting for the data team, a new model to try. A year later the company has several pilots and no production AI.

02Why do pilots stall?

The reasons are rarely about model quality. They are about what the pilot was never designed to answer.

Success was never defined. “Users liked it” is not a decision criterion. Without a measurable outcome (time saved per ticket, extraction accuracy, resolution rate) and a baseline, there is no point at which the pilot has clearly succeeded, so it never clearly ends.

There is no eval set. Quality was judged by people trying it and forming impressions. When the model or prompt changes, nobody can say whether it got better or worse, so nobody is confident enough to put it in front of customers (see ship a prompt change with an eval).

It ran on a copy of reality. The pilot used an exported spreadsheet, a sample of documents or a sandbox with no permissions. Connecting to live systems, with access control and current data, turns out to be most of the work, and it was all left for later (see out of context).

It lived beside the workflow. Users had to open a separate tool to try it. Usage looked fine in the pilot, when people were asked to use it, and would have collapsed in daily work (see put the model inside the workflow).

Risk was deferred. Data protection, security, audit and failure handling were “for production”. When production approaches, those questions arrive all at once and the project stops (see data compliance and audit trails).

Nobody owns it afterwards. The pilot was run by an innovation team or an outside vendor. No product or engineering team has agreed to operate it, on-call, for years.

03What does a production-minded pilot look like?

Design the pilot as the first release of a production system, at a small scope.

Instead ofDo this
A broad demo across many use casesOne workflow, end to end
“Users liked it”A target metric, a baseline and a threshold agreed upfront
Ad hoc testingAn eval set of real cases, run on every change
Exported sample dataLive connectors with real permissions, even if read-only at first
A separate pilot appA feature inside the tool people already use
Risk review at the endData flows, logging and failure handling designed in week one
An innovation team that hands overThe team that will run it, involved from day one
pilots/claims-triage/charter.yaml
workflow: insurance-claims-triage
owner: claims-platform-team            # operates it after the pilot
outcome: { metric: minutes_to_first_action, baseline: 42, target: 15 }
quality: { eval: evals/claims_triage.jsonl, bar: { routing_accuracy: 0.93 } }
scope: { lines: [motor], regions: [IN], data_classes: [policyholder_contact, claim] }
integration: { reads: [claims_db, policy_admin], writes: [triage_queue], ui: claims_workbench }
risk: { gateway: required, pii_redaction: on, human_review: amount > 200000 }
decision_date: "+10 weeks"             # ship, extend scope, or stop

A charter like this makes the end of the pilot a decision, not a drift. On the decision date the numbers say ship, extend, or stop, and each of those is a success compared with purgatory.

A pilot should answer one question: would we run this in production? Design it so the answer is a number.

04Where does this fit?

Pilot purgatory is an L5 product and workflow problem with L2 roots: no evals, no measured quality. It is why we build engagements as production systems at small scope from the first week, on the foundations described in the owned intelligence layer.

Previous: Avoid LLM vendor lock-in. Next: No audit trail.

Building something like this?DESCRIBE A SYSTEM →