AI industry

The 95% is not a model problem

A sleek device glowing faintly, still sealed in its shipping box on an aluminum table — bought, never deployed

MIT's NANDA initiative studied enterprise generative-AI deployments and found that roughly 95% of pilots delivered no measurable impact on the P&L. About one in twenty produced real acceleration. Everything else stalled somewhere between the demo and the operation.

The number got passed around as proof that AI is overhyped. That's the wrong reading, and it's a comfortable one — it lets everybody off the hook. The researchers were specific about the cause: the failures weren't about model quality. They were about the learning gap. General-purpose tools are brilliant for an individual because they're flexible, and they stall inside a company because they don't learn the company — not its procedures, not its terminology, not the way its work actually moves.

I'd put it in plainer terms. A general model knows your industry. It doesn't know your business. And the question that blocks a real person on a real Tuesday isn't "what is best practice for this?" — it's "how do we handle this here?" A tool that can't answer the second question is a very expensive search engine.

The same study found something more damning than the headline: a majority of firms evaluated enterprise-grade systems, a fraction piloted them, and only a sliver went live. Investment isn't the constraint — billions went in. Integration is. One manufacturing COO put it to the researchers better than any consultant could: the hype says everything has changed, but in their operations, nothing fundamental had shifted.

Here's what I think the 5% do differently, and it isn't glamorous.

They fix the process before they point intelligence at it. Automation multiplies whatever it touches. Point it at a clean workflow and you get speed; point it at a broken one and you get mistakes at scale, delivered with confidence. Most pilots skip the unglamorous mapping phase because it doesn't demo well.

They ground the system in their own reality. Their procedures, their history, their vocabulary, their exceptions — wired in, not prompted at. That's the difference between a tool that answers and a tool that answers correctly for you.

They name a number before they start. A KPI and a Week-0 baseline, agreed in advance. Pilots without a baseline can't succeed, because success was never defined — they can only produce enthusiasm, and enthusiasm has a shelf life of about a quarter.

They keep a human in the loop where it counts, which is what makes the thing safe enough to actually deploy rather than perpetually "almost ready."

None of that is an argument against AI. It's an argument against skipping the foundation and calling the resulting rubble a technology failure. The models are extraordinary and getting better weekly. The operations they're being dropped into are mapped in nobody's head, documented nowhere, and owned at the seams by no one.

The 95% didn't fail because the intelligence wasn't good enough. They failed because it was pointed at something nobody had bothered to understand first.

Sources: Fortune on MIT NANDA, The GenAI Divide: State of AI in Business 2025 · Forbes analysis of the same report

Map it before you build it