There's a statistic floating around that some large share of enterprise AI pilots never reach production. I don't know if the exact number is right, but the shape of it matches what I see. The pilots mostly work. The demos go well. Everyone agrees it's promising. And then nothing happens. Six months later the pilot is still a pilot, the sponsor has moved on, and the team has stopped mentioning it in stand-ups.
What's frustrating is that the reasons are rarely technical. I've been called in to rescue enough of these to have a list, and the list is boring. That's the good news, because boring problems have boring fixes.
Five ways pilots die
1. Nobody owns it after the demo
A pilot usually has a champion: someone who wanted it, found budget for it and showed it to the executive team. A production system needs an owner, which is a different job. The owner keeps it running, watches its quality, fields the complaints and argues for its budget next year.
The handover from champion to owner almost never happens by itself. The best predictor I've found for whether a pilot ships is whether anyone can answer "who gets paged when it's wrong?" If the answer is a shrug, the pilot is already dead. It just doesn't know yet.
2. The pilot ran on clean data
Pilots get the curated dataset. Fifty hand-picked documents, a tidy export, a test tenant. Production gets the SharePoint site with three contradictory versions of every policy, the CRM where half the fields are free-text dumping grounds, and a ticket history full of "see attached" with nothing attached.
The fix is unsatisfying but it works. Run at least part of the pilot on the worst data you have. If the system holds up against your messiest content, production is an expansion rather than a surprise. If it doesn't, you found out for the price of a pilot instead of the price of a launch.
3. Security and legal saw it last
A team builds for three months, books the sign-off, and then the security review asks questions that should have been design inputs. Where does the data go? What does the vendor retain? What's the plan for prompt injection? Who checked the permissions on that connector? Every question is fair, and every answer now means rework. The momentum dies in the queue.
Flipping the order is cheap. Spend an hour with your security team in week one and ask them what would make them reject the thing you're about to build. The review becomes a checklist instead of a gate. I've watched that one meeting take months off a delivery timeline.
4. Nobody defined "good enough"
Pilots get judged on impressions. People try it, it seems clever, thumbs up. Production needs a number. What accuracy, on which test set? How many escalations to a human are acceptable? At what error rate do we switch it off?
Without an agreed threshold, every mistake turns into a vote on the whole system. One bad answer lands in the wrong inbox and suddenly the project is "unreliable", even if it's right 96% of the time and the process it replaced managed 89%. Set the bar before launch, measure against it and publish the results. Systems with published numbers survive their first bad week. Systems judged on vibes don't.
5. The cost model was a guess
The pilot cost a few hundred dollars a month, so nobody did the maths for ten thousand users. Eventually someone does, the number has more digits than anyone expected, and the business case has to be argued again from scratch in front of a sceptical audience.
Token costs in production are mostly an architecture problem. Caching, routing and output limits can change the bill several times over. That work belongs in the pilot, while the architecture is still cheap to change, not after the first big invoice.
What shipping teams do differently
The teams I've seen get pilots into production have a few habits in common, and none of them are glamorous.
They build the pilot inside production constraints from the start: real auth, real data permissions, real network rules. It's slower to get going and much faster to finish.
They name the owner on day one. I mean the person who'll run it, not the champion. That person tends to shape the pilot into something they can operate.
They write the launch criteria before the pilot begins, usually three or four measurable conditions. When those are met, it ships, and nobody gets to move the goalposts in either direction.
And they keep the scope embarrassingly small. One workflow, one group of users, one integration. Pilots that try to prove everything tend to prove nothing in time to matter.
Ask this before you start
Pilots almost always work, so "will this work?" isn't a very useful question. Ask "what would stop this from shipping?" instead. Ask it in week one, write the answers down, and spend the pilot getting rid of those risks rather than polishing the demo.
Got a pilot that worked and stalled anyway? I'm happy to help you figure out which of these it ran into.

