FIG. 02
Four reasons AI pilots never reach production
The demonstration went well. The management meeting agreed it was promising.
Six months later nothing is in production, and nobody can say why.
There are four causes. It is usually some combination of them, and rarely anything more exotic.
1. It only ever saw the happy path.
The prototype was built over a fortnight and tested on the documents somebody chose to test it on: the clean ones in the expected format, mostly from the three largest suppliers. Nobody fed it the scanned fax, the invoice with two purchase order numbers, or the credit note that arrives once a quarter and breaks every assumption in the process. These get filed as edge cases. In back-office work they are a routine part of the volume, and they are where the engineering actually is.
2. Nothing was measured beforehand.
Nobody wrote down what the process cost before the pilot started. So when the pilot finished, there was no way to settle the argument about whether it had helped. One person said it was faster. Another said the checking took as long as the typing used to. Both were reasoning from impressions, because impressions were the only evidence anyone had collected.
A pilot with no baseline can’t be defended in the budget meeting. It can only be believed or disbelieved, and by then the room has moved on.
3. It was never integrated.
It lived in a browser tab next to the ERP. Somebody read the output on one screen and typed it into the other. The work moved rather than disappeared, and the second window is a cost in itself. A system that isn’t inside the process the company already operates is a tool people have to remember to use, which means they stop.
4. Nobody owned it on the Monday after go-live.
This is the one nobody plans for. The developer who built it moved to the next project, and the model provider deprecated the version it was written against. The output quality drifted and no one was watching, because nobody had been made responsible for watching. Nothing failed loudly. It just got worse until people quietly went back to doing it by hand.
None of these four is an AI problem. Three of them are project problems and one is an operating problem. Which is why picking a better model doesn’t fix any of them.
The version that works is unglamorous: measure the process for a week first, so the result can be argued from evidence instead of impressions; build for the irregular inputs rather than the demonstration; put it inside the system people already have open; and name in writing, before go-live, who is accountable for it in month 14.
If your pilot died, it was one of these four. Usually the fourth.