Perspective

Artificial Intelligence and Data

Why most AI pilots never leave the pilot

A pilot proves a model can work. It does not prove an organisation can run one.

Daniel Okonkwo · Data and AI · 3 June 2026 · 5 min read

A close-up of a circuit board with a processor at its centre

Only about one pilot in eight becomes something people use on an ordinary Tuesday. The gap is not explained by model quality. Most stalled pilots worked fine in the demo, which is exactly why the stall is so frustrating to everyone involved.

What a pilot demonstrates is feasibility under favourable conditions: a curated dataset, a motivated team, and no obligation to handle the case that only shows up twice a month.

The three questions a pilot never has to answer

Who fixes it at two in the morning? What happens when the model is confident and wrong? And who absorbs the cost of the change to how people work?

None of these are research questions, and none of them get easier by improving the model. They get answered by deciding, early, that the pilot is a step toward an operational service rather than an exhibit.

Pick the use case for the second question

Teams choose pilots for visibility. Executive attention flows to whatever sounds most transformative, which is usually the use case with the widest blast radius and the least tolerance for error.

The use cases that survive tend to be narrower and duller: a step in an existing process where the output is checked by someone who would have done the work anyway, and where being wrong costs a minute rather than a customer.

Measure the workflow, not the model

Accuracy on a held-out set tells you very little about whether anything improved. Time to resolution, rework rate, and how often people override the system tell you a great deal.

The override rate in particular is worth watching. When it drifts up, something changed upstream, and it is usually the data.