Do 95% of AI projects actually fail?
No.
The number is real in the sense that somebody published it. It is not real in the sense that anybody measured it. The study it comes from says something considerably narrower, and the narrower thing is more interesting than the headline.
I went looking because I was about to cite it myself.
Where the number comes from
The source is a July 2025 report from MIT’s Project NANDA, The GenAI Divide: State of AI in Business 2025. Its executive summary says 95% of organizations are getting zero return on their generative-AI investment. That sentence went around the world.
The report’s body does not support it.
What the body actually contains is a funnel for custom, task-specific enterprise GenAI tools: 60% investigated, 20% piloted, 5% reaching what the authors call successful implementation. That 5% is where the 95% comes from — it is the complement of one number in one category.
In the same chart, general-purpose tools like ChatGPT and Copilot run 80% investigated, 50% piloted, and 40% successfully implemented. Same report, same exhibit, eight times the success rate. The headline picked the worse of the two and dropped the qualifier.
The methodology is 52 structured interviews, 153 survey responses collected at four industry conferences, and a review of 300-plus publicly announced initiatives, between January and June 2025. The report is candid about what that supports:
These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category.
That is an honest caveat, and it is doing a lot of work. “Directionally accurate, based on interviews” is not a measurement. It is a well-informed impression, which is a legitimate thing to publish and a poor thing to quote as a statistic.
Kevin Werbach at Wharton put the problem more bluntly than I would have:
If MIT Project NANDA stands behind the claims, it should release the full supporting data. If not, it should retract the report.
Verdict: contested, and rejected as usually stated. There is no methodology behind “95% of AI pilots fail”. There is a 5% figure for one subcategory, sitting next to a 40% figure for another, in a preliminary report that says its own numbers are directional.
The second number
“Gartner predicts 30% of generative-AI projects will be abandoned after proof of concept by the end of 2025.”
This one is quoted as though the year has been graded. It has not. The statement is a forecast, published in July 2024, about a year that had not happened yet.
Rita Sallam, the analyst quoted in it, was describing pressure rather than outcomes:
After last year’s hype, executives are impatient to see returns on GenAI investments, yet organizations are struggling to prove and realize value.
I could not find any retrospective — from Gartner or anyone else — establishing what the actual 2025 abandonment rate turned out to be. So the honest status is: a reasonable forecast, made by people with good visibility, that nobody has checked.
Worth noting alongside it: Gartner’s June 2025 prediction that over 40% of agentic AI projects will be cancelled by end of 2027 names the same three causes as the 2024 one — escalating costs, unclear business value, inadequate risk controls. Identical diagnostic categories, different technology wave, a year apart. That tells you something about the categories.
Verdict: plausible as a forecast, unverified as an outcome. Cite it as a prediction or not at all.
The third number, which is the instructive one
“47% of enterprise AI users have based a major business decision on hallucinated content.” Usually attributed to Deloitte. It travels with two companions: $67.4 billion in losses in 2024, and $14,200 per employee.
I tried to find it. The trail runs from one blog to another blog to a citation reading “Deloitte, 2025” — no report name, no page, no methodology.
So I went to Deloitte’s actual State of AI in the Enterprise report, the 2026 edition, 3,235 senior leaders across 24 countries, the largest and most methodologically transparent survey in any of this. The figure is not in it. Neither are its two companions.
Verdict: rejected. Not misinterpreted, not out of date — as far as I can tell it does not come from anywhere. It looks Deloitte-sourced because a chain of posts said so, and each link in that chain was written by someone who trusted the previous one.
Two more, quickly, from the technical side
“Lost in the middle.” The finding that a model degrades on information buried in the middle of its context. It is the most-cited claim in context engineering, and it is real — for the models it was measured on in 2023: GPT-3.5-Turbo, Claude 1.3, MPT-30B, LongChat-13B. All obsolete.
A replication published in May 2026 re-ran it on current open models and found the classic U-shaped curve mostly does not reproduce; accuracy was, in their words, comparatively flat across positions. So the field is citing a 2023 result about 2023 models as a description of how today’s models behave.
“Context rot starts at 300–400K tokens.” The underlying research is real and useful: Chroma tested 18 current models in July 2025 and found reliability degrading with input length. The specific threshold numbers circulating in secondary coverage are not in their report. Someone invented a round figure and it propagated.
Chroma also sells vector databases, and “long context is not a free lunch, you still need good retrieval” is exactly what a retrieval vendor would want to be true. That does not make the data wrong — it is multi-model and methodologically clear — but the funder belongs next to the finding, every time.
Why this keeps happening
None of these numbers spread because people are credulous. They spread because each one confirms something the reader already believes.
Everyone in an enterprise has watched an AI initiative underdeliver. “95% fail” arrives as vindication, and vindication does not get audited. The claims that survive unchecked are precisely the ones nobody wants to check, which is the opposite of how it should work.
The tell is structural, and you can apply it without leaving your desk. A statistic that cannot name its population, its n, and its measurement is not a statistic. All three of the failed claims above collapse on the third one. Nobody measured whether an AI pilot “failed”; they asked people at conferences how it was going.
What actually survives
This is the part that matters, because the phenomenon is real even though the evidence for it is not what people think.
Adoption is broad and impact is shallow. McKinsey’s State of AI survey, published November 2025, n=1,993 across 105 countries: 88% of organizations use AI in at least one function, while 39% can point to any enterprise-level EBIT impact — and most of those put it below 5%. Self-reported, like nearly all survey data here, but consistent across many independent reports.
The model itself does work. This is the finding that gets lost. Cui and colleagues ran genuine randomized controlled trials at Microsoft, Accenture and a Fortune 100 company — 4,867 developers, randomly assigned an AI coding assistant, 26% more tasks completed. Measured output, not self-report, published in Management Science.
So the shape of the real problem is not “AI does not work.” It is that a measured 26% individual gain does not show up in the enterprise’s accounts. Something between the individual and the P&L is absorbing it.
The MIT report is actually good on this, in a section that got none of the attention its headline did:
The biggest thing holding back AI is model quality, legal, data, risk → What’s really holding it back is that most AI tools don’t learn and don’t integrate well into workflows.
Don’t learn is the phrase to sit with. A tool that cannot retain what it was told last week, cannot be corrected once, and cannot carry an organization’s own context into the next task will produce exactly this pattern: real gains in the moment, nothing that compounds.
The honest limit
I would like to tell you this proves the bottleneck is organizational rather than technical. It does not, and the pieces claiming otherwise are overreaching.
Every source above is either a survey asking people to attribute blame — which is the kind of question respondents are worst at answering — or a benchmark that says nothing about organizations, or an RCT that measures individuals rather than balance sheets. No controlled study isolates the two variables.
And there is real counter-evidence. METR tracks how long a task can be before frontier agents stop completing it reliably. Agents succeed on tasks of a few minutes almost always, and on tasks beyond roughly four hours less than 10% of the time. That horizon is doubling every seven months or so, but it is doubling from a low base — and “complex, multi-week workflow” describes most of what an enterprise actually wants automated.
So: capability limits are real at the long-horizon end. The organizational explanation is the dominant reading of 2025–2026, and it is well-argued, but it is a reading rather than a demonstration. I will take it as far as the evidence goes and no further.
What to do with this
Ask three questions of the next AI statistic you are handed. Who was measured, how many of them, and what instrument was used. If the answer to the third is “we asked them,” you have an impression, and impressions are fine as long as nobody calls them findings.
I keep a register of every number I use, with its source and its grade, because I have now been wrong about one of them in public. That habit is not scrupulousness. It is the cheapest available insurance against building an argument on a figure that turns out to be a chain of blog posts citing each other.