Evidence

Why the 95% AI failure stat is real, and what's actually behind it

MIT’s 2025 review of enterprise AI initiatives put the failure rate at 95%. RAND landed somewhere above 80%. These numbers get cited in boardrooms and conference talks, usually to make a point about how hard AI adoption is, and then the conversation moves on to tools, platforms, vendors.

Which is exactly the problem.

Nobody’s asking what’s causing it. The assumption, usually unspoken, is that the technology is difficult, immature, or being applied wrong. That the right model, the right vendor, or the right implementation partner would move the needle. So organisations procure differently, implement more carefully, hire better consultants — and the failure rate stays roughly where it is.

It’s not a technology problem

I’ve spent enough time inside organisations running AI initiatives to have a reasonably clear picture of what’s actually going wrong. The technology is rarely it.

What I find, almost without exception, is some combination of the following. Nobody has clearly defined which decisions the AI is supposed to improve, or how you’d know if it had. The data feeding the model is messier than anyone admitted at the start, because the teams who own it have never had a reason to clean it up. The governance around AI outputs is either missing or theoretical — someone wrote a policy, nobody reads it, and there’s no clear answer to the question of who’s responsible when the model is wrong. And leadership is announcing AI adoption while doing nothing to resolve the organisational conditions that will determine whether it works.

None of these things appear on a technology roadmap. They don’t come up in vendor demos. They’re the kind of problem that only becomes visible when you look at the organisation rather than the tool.

Why the measurement problem matters

Part of the reason the failure rate is so high is that failure is hard to see in real time. AI initiatives don’t usually die dramatically. They produce results that can’t be attributed to anything specific. They get quietly descoped. They get defined as a success because a pilot went well, with no follow-up on whether it scaled or changed anything.

If you didn’t define success before you started, you can’t measure whether you got there. That sounds obvious. It happens constantly.

The organisations that do get something real from AI tend to have done three things before anyone wrote a line of code. They agreed what they were trying to change and how they’d know it had changed. They audited the data situation honestly, including the parts that were embarrassing. And they were clear about who owned the AI output when it was wrong — not in a liability sense, but in a practical one.

What the stat is actually measuring

The 95% failure rate is not measuring whether AI is useful. It’s measuring whether organisations are ready to use it. Those are different questions, and most of the industry is treating them as the same one.

The tools are not the constraint. The system those tools sit on top of is. Until that changes, the failure rate won’t.