The AI pilot that never made it to production
Five reasons it happens, and which one is usually yours.
A pilot proves a model can do something. Production asks who owns it, what happens when it is wrong, where the data comes from every day, and whose week changes. Most pilots die because nobody answered those four before building, not because the model was bad.
Work it out in a minute
Six questions on one task, scored on volume, rules, stability and judgment. Returns a straight verdict.
Where pilots die
Four gates between a demo and production. The model is on none of them.
Demo
A baseline
what it cost before
A data path
arriving daily, from a system
An exception route
who sees the wrong ones
A named owner
whose job it is afterwards
Production
What you get
It was measured against a demo, not a baseline
Nobody wrote down what the task cost before. So when the pilot works, there is no number to compare it with, no business case to fund the next step, and the decision falls back on whoever is most enthusiastic.
The data was hand-fed
Somebody exported a clean file for the pilot. In production that file has to arrive every day, from a system, with the same columns, when somebody is on holiday. That pipeline is usually more work than the model was.
Nobody owned the wrong answers
The pilot had a human checking everything, which is how it looked accurate. Production needs a rule for which cases go to a person, how they are told, and what the system does meanwhile. Without it, it cannot be switched on.
Nobody owned it afterwards
The person who championed it went back to their real role. A system with no named owner does not fail loudly. It drifts, then quietly stops being trusted, then stops being used.
It needed a process change nobody agreed to
The model works, and using it means a team changes how they work. That conversation never happened, because the pilot was run as a technology project rather than an operations one.
The scope grew while it was being proved
It started as one task. By the end of the pilot it was three departments and an integration, which is now too expensive to approve and too big to start. Scope creep kills more pilots than accuracy does.
How to tell which one killed yours
Ask what the number was before. If nobody can answer, it was the first one, and it is the most common by a distance. A pilot without a baseline cannot be approved, because there is nothing to approve it against.
None of these are model problems, which is the point. They are operations problems wearing a technology costume, and they are the reason a pilot can be a technical success and a commercial dead end at the same time.
Questions we get
Was the model the problem?
Usually not. Most stalled pilots had accuracy good enough for the job and failed on ownership, data supply or exception handling.
If accuracy really was the issue, that is normally visible in the first fortnight, not at the production gate.
Can we restart without throwing it away?
Nearly always. The prompts, the evaluation set and whatever was learned about the data keep their value. What gets rebuilt is the part around them.
Start by writing down the baseline that was never taken.
How long does it take to get it live?
When the model already works, the remaining work is a data path, an exception route, a named owner and a measurement. That is typically weeks.
The long pole is usually agreeing the process change, not building anything.
Who should own it once it is live?
Somebody in the operation it serves, not in the technology function. They are the person who notices first when the outputs stop making sense.
Give them a number to watch and a route to report it.
The same four gates, ahead of time, are in preparing for AI. On who picks it up from here, an AI engineer or a consultant. And if that means bringing somebody in, six questions to ask them.
More in the guides and every answer in one place.
Read next
What to settle before the next one starts, the rules it runs under, and who should be doing the building.
Shaheer leads the work, with engineers, writers, filers and analysts behind him. C-suite operations for a San Francisco AI company, Six Sigma on the process side, Anthropic certified on the Model Context Protocol, ten years across eight industries. See what we have built