Why so many AI pilots never reach production
The failure almost never happens in the model. It happens at the handover, when a pilot owned by innovation has to become a system owned by operations, and nobody funded that transition.
Every company I have worked with can point to a pilot that worked. The model hit its accuracy target, the demo landed, the executive sponsor was pleased. Two quarters later the same pilot is a dormant repository and a slide in a deck about learnings. Ask what went wrong and you will hear about data quality, or vendor limitations, or shifting priorities. Those are symptoms. The cause is almost always structural.
No single figure captures how often this happens, and the ones in circulation measure different things. RAND's 2024 study of AI project failure, based on interviews with 65 data scientists and engineers, reports estimates that more than 80 percent of AI projects fail — roughly twice the failure rate of IT projects that do not involve AI. IDC's AI CIO Playbook 2025 counts conversion instead: for every 33 proofs of concept an organisation starts, about four reach production. MIT's 2025 GenAI Divide report counts return, finding that around 95 percent of the generative AI pilots it examined showed no measurable financial impact. Abandonment, conversion and return are three different questions with three different answers. What the research agrees on is the direction, and on where the cause sits: organisational readiness, data and deployment infrastructure rather than model capability.
The handover nobody funds
A pilot is funded as an experiment. It has a small budget, a tolerant audience, and permission to be rough. Production is funded as infrastructure. It needs monitoring, an on-call rotation, an access model, a retraining cadence, a rollback plan, and a named owner who is accountable when it misbehaves at two in the morning. Those are two different budget lines, two different approval paths, and usually two different departments.
The pilot team is not staffed to build the second thing, and the operations team was not consulted about the first. So at the moment the pilot succeeds, it arrives at a door nobody has the key to. Innovation cannot run it forever, operations will not accept it without the scaffolding, and the sponsor who funded the experiment has no line item for the difference.
A pilot that cannot name its production owner on day one is not a pilot. It is a demonstration.
What the receiving team actually needs
When I ask operations leaders what would make them willing to take on an AI system, the answers are consistent and unglamorous. They want to know how the system fails and how they will find out. They want a threshold that tells them when to stop trusting it. They want the ability to turn it off without turning off the business process it sits inside. They want to know who fixes it and how fast.
None of that is a modelling problem. All of it is a design problem, and all of it is cheaper to solve at the start of the pilot than after it succeeds. A pilot that ships with a monitoring plan, a documented failure mode, and a manual fallback path costs perhaps fifteen percent more to build and is several times more likely to survive the handover.
Three things that change the odds
Name the production owner before the first sprint. Not the sponsor, not the innovation lead. The person whose team will run this in eighteen months. Give them veto power over the design. If nobody will accept that role, the use case is not ready and the pilot should not start.
Budget the transition as part of the pilot. A single number covering both phases removes the second approval, which is where most pilots die. If the combined cost cannot be justified, you have learned something valuable before spending anything.
Pick the boring use case first. The most durable AI system in most companies is the one that removes a queue nobody defends. It has a clear owner, a measurable baseline, and no political shadow. Save the ambitious use case for after you have proven the pipeline from experiment to operations works at all.
The organisations that make it through the handover are rarely the ones with better models. They are the ones that treated the handover as the project and the model as a component of it.
- James Ryseff, Brandon F. De Bruhl and Sydne J. Newberry, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, RAND Corporation, RR-A2680-1, 2024. rand.org
- IDC, The AI CIO Playbook 2025 (commissioned by Lenovo) — 33 proofs of concept per four production deployments.
- MIT NANDA initiative, The GenAI Divide: State of AI in Business 2025 — 95 percent of surveyed generative AI pilots showing no measurable P&L return.
- —Pilots fail at the handover, not the model. Experiment budgets do not cover production scaffolding.
- —Name the production owner before the first sprint and give them design veto.
- —Fund pilot and transition as one number so there is no second approval to lose.
- —Start with the queue nobody defends, not the use case with the biggest headline.