← INSIGHTS

Why AI pilots fail, and what the ones that survive do differently

The "95% of AI pilots fail" statistic gets misread constantly. Here's what MIT's report actually found, and what we see separating pilots that survive from ones that quietly fade.

BY BETTINA MEYER··7 MIN READ

The statistic everyone quotes, and what it actually says

If you've sat in a leadership meeting about AI in the past year, someone has probably said "MIT found that 95% of AI pilots fail." It's become the default bear case, cited in board decks and consultant slides on every continent.

Here's what the report, MIT NANDA's GenAI Divide: State of AI in Business 2025, actually measured. Researchers interviewed 150 executives, surveyed 350 employees, and analysed 300 AI projects. Of those 300 projects, 95% delivered no measurable impact on profit and loss.

That distinction matters, because if you build your AI strategy around avoiding "failure," you'll spend your energy on the wrong thing. If you build it around designing for measurable impact and real adoption, you're closer to what's actually going on.

The wider data on AI projects tells a similar, if slightly more encouraging, story. IDC research commissioned by Lenovo, reported in March 2025, found that 88% of AI proof-of-concepts never made it to production, roughly four out of every 33 launched. The same research programme's follow-up study in January 2026 found that figure had improved to 46% of proof-of-concepts reaching production, alongside 60% of organisations now in late-stage AI adoption. The direction is positive. It's also worth noting that only 27% of those same organisations report having a comprehensive AI governance framework in place, which is its own kind of readiness gap.

The finding worth acting on: buy tends to beat build

Buried under the headline is a more useful number. Purchasing AI tools from vendors and partnering with them succeeded about 67% of the time in the projects studied. Building the equivalent capability in-house succeeded at roughly a third of that rate. For a business scoping its first serious AI initiative, that's a real signal: unless there's a specific, defensible reason your internal build will outperform that base rate by a wide margin, buying and integrating an existing tool is the safer starting point.

This tracks with something we see constantly. A leadership team convinced that AI is "core to the business" often assumes that means building proprietary tooling. In practice, the differentiator is rarely the model. It's how well the tool is embedded into the specific workflow, data, and decisions of that business, and that's exactly where a vendor with deployment experience across many customers usually has the edge over a team building from scratch.

The pattern holds at the macro level too. Harvard Business Review reported in September 2026 that BCG's January 2026 AI Radar found corporations expect to roughly double AI spending this year, yet McKinsey QuantumBlack's April 2026 analysis found 60% of organisations still see no enterprise-wide EBIT impact, even as nearly eight in ten use generative AI in at least one business function. The authors' conclusion is blunt: companies are automating individual tasks rather than redesigning the workflows that actually produce business value, which is a more precise way of stating the pattern we see in pilots that stall.

What we see: pilots don't die of bad technology, they die of no ownership

Here's the pattern we've watched repeat across pilots that quietly fade rather than formally fail. Someone runs a proof of concept. It works well enough in the demo. Leadership is pleased. Then the pilot moves from "interesting experiment" to "thing a busy operations manager is now supposed to maintain on top of their actual job," and nobody has been given the time, authority, or incentive to make that happen. Three months later the licence is still active, the dashboard shows a handful of logins, and everyone has quietly moved on without ever calling it a failure.

We'd estimate, from what we've observed across engagements, that this pattern accounts for more dead pilots than any technical shortfall in the tool itself. A senior AI leader who has watched a dozen of these die learns to ask a blunt question before a pilot even starts: who, specifically, is accountable for this still being used in six months, and have they been given the authority to change the workflow around it, not just the tool inside it. If the honest answer is "the person who happened to run the pilot, in their spare time," that's the failure mode playing out before day one.

The second common cause is scope. Enterprises in the MIT data skewed toward deploying generative AI in marketing and sales, where it's easier to greenlight a pilot, while the report found more measurable value sitting in back-office processes: the two-day-a-month reporting cycle, the manual data entry between two systems, the recruitment step nobody has automated because it's unglamorous. Marketing use cases are visible and easy to demo. Back-office use cases are where the cost actually gets removed, and removed cost is far easier to measure than a marginal lift in campaign performance.

What the pilots that survive do differently

The pilots that make it past the demo stage share a few habits, drawn from what we've seen work.

They start with a workflow, not a tool. Instead of "let's try this AI product," the starting question is "this specific process is slow or expensive, what's the smallest change that would fix it." The tool follows from the problem, not the other way round.

They name an owner before launch, not after. Someone with the authority to change how a team works, not just administer a licence, is accountable for adoption, and that person's job changes to reflect it.

They measure before they deploy. Even a rough baseline, how long the task takes now, how much it costs, how often it goes wrong, is enough to tell six months later whether anything actually improved, rather than relying on a gut feeling that people seem to like it.

They buy before they build, unless there's a clear reason not to. Given the gap in the data between purchased and in-house success rates, a team defaulting to a custom build should be able to explain specifically why their situation is the exception.

They expect to redesign the workflow, not just insert the tool into the existing one. The report's lead author pointed to startups picking a single pain point and executing well as a factor behind their stronger results, rather than any advantage in the technology itself. An established business inherits a workflow built for a world without AI in it, and slotting a tool into that unchanged process is a common way pilots quietly underperform.

Where this points next

None of this is really about the technology being immature. It's about treating a pilot as a real change to how work gets done, with an owner, a baseline, and a workflow redesigned around it, rather than a tool switched on and hoped for. That's closer to an AI strategy and transformation programme than a procurement decision, which is why the businesses getting real value tend to treat it that way from the start.

FAQ

Did MIT really find that 95% of AI pilots fail?
Not quite. The report found that 95% of the 300 AI projects studied showed no measurable profit and loss impact, which is a narrower and different claim than "the pilot failed." A tool can be genuinely useful and still show no measurable financial impact if nobody set a baseline to measure against.

Is it better to build AI tools in-house or buy them?
In the projects MIT studied, buying and partnering with vendors succeeded about 67% of the time, versus roughly a third of that rate for internal builds. That doesn't rule out building your own, but it sets a bar: you should be able to explain why your situation beats that base rate before committing to it.

Why do AI pilots stall after a promising start?
Most often because no one was given clear ownership and authority to change the surrounding workflow once the pilot moved past the demo stage. The tool keeps running, but nobody's job changed to make sure it kept getting used.

Should we start with marketing and sales, or back-office processes?
Back-office processes tend to produce more measurable value, because cost removed from a process is easier to track than a lift in campaign performance. Marketing pilots are common because they're easy to greenlight, not because they deliver the most value.

How do we know if our pilot is actually working?
Set a baseline before you deploy: how long the task takes, what it costs, how often it goes wrong. Without that, six months later you're relying on impressions rather than evidence either way.

WORK WITH PRAXES

Ready to move from reading about AI to getting your organisation ready for it? Start with the diagnostic.

Take the AI Readiness Diagnostic +