The demo goes well. Everyone in the room can see it working. Six months later it is still a pilot, the enthusiasm has drained, and the honest answer to “what happened?” is that nothing did.
This pattern is common enough to be predictable, and the cause is almost never the model.
The pilot was scoped from the tool, not the constraint
Most stalled pilots start with a capability rather than a problem. A vendor demonstrates something impressive, somebody asks where it could be applied, and a use case is worked backwards from the demo.
Pilots scoped this way fail a specific test: nobody can say what would change if it worked. Not what it would do — what would change. Fewer hours in a process. Faster close. Fewer escalations. Without that, there is no threshold for success, so the pilot cannot pass and cannot be killed. It just persists.
The productive version of the question starts at the other end: where does expensive human judgement get applied repeatedly to similar inputs? That is where a model has a chance of changing the economics.
The data was assumed rather than checked
The second failure arrives about six weeks in, when someone tries to assemble the training or retrieval data and discovers that it is spread across three systems, that two of them disagree, and that the field everyone assumed was populated is populated about forty per cent of the time.
This is discoverable in a day, before anyone commits to a pilot. Ask for a sample of the actual data — not a description of it, a sample — and check three things: is it complete enough, is it consistent enough, and can a system reach it in the way the workflow would need to.
There was no path from pilot to production
A pilot is not a small version of a production system. Production requires things pilots skip: identity and access, audit logging, monitoring, a support path when it is wrong, a rollback, and someone whose job includes owning it.
If nobody has scoped that work, the pilot will succeed and still not ship — because the distance between “it works” and “it runs” turns out to be most of the project, and it usually needs a budget that was never requested.
Scope the production path at the same time as the pilot, even roughly. If the answer is that production would take three quarters and a headcount, that is worth knowing before the pilot, not after.
The people affected found out last
The final failure is human. A pilot that changes how a team works, designed without that team, arrives as something being done to them. What follows is not sabotage; it is something more corrosive — polite non-adoption. The system is available. People keep doing it the old way. Usage numbers stay flat and nobody can quite explain why.
The teams that get past this involve the people doing the work in defining what good looks like, and are explicit about what happens to their role. Ambiguity on that point is read, correctly, as bad news.
The four conditions
Before starting an AI pilot, four things should be true. If any is missing, fix it first — the pilot will not fix it for you.
- A named business metric that would move, with a number attached and a person who cares about it.
- Data that has been looked at, not described — sampled, checked for completeness and consistency, and confirmed reachable.
- A scoped path to production, including who owns it, what it costs, and how it gets turned off if it misbehaves.
- The affected team in the room from the start, with an honest account of what changes for them.
None of this is about the technology being hard. The models are the easiest part of the problem now. What has not gotten easier is the organisational work of deciding what is worth automating, checking whether the data supports it, and bringing the people who do the work along with you.
That is the work that determines whether the pilot ships.