# Do 95% of AI projects actually fail?

CM-001 · written 2026-08-19 · published 2026-08-19 · guide
tags: ai-adoption, evidence, context-management

> No. The study everyone cites says something much narrower, and four of the field's other most-repeated statistics do not survive contact with their sources either.

No.

The number is real in the sense that somebody published it. It is not real in the sense that
anybody measured it. The study it comes from says something considerably narrower, and the narrower
thing is more interesting than the headline.

I went looking because I was about to cite it myself.

## Where the number comes from

The source is a July 2025 report from MIT's Project NANDA, *The GenAI Divide: State of AI in
Business 2025*. Its executive summary says 95% of organizations are getting zero return on their
generative-AI investment. That sentence went around the world.

<aside class="margin-note">
MIT Project NANDA, <em>The GenAI Divide: State of AI in Business 2025</em>, July 2025. The report is
self-labelled "Preliminary Findings" and is not peer-reviewed.
</aside>

The report's body does not support it.

What the body actually contains is a funnel for **custom, task-specific enterprise GenAI tools**:
60% investigated, 20% piloted, 5% reaching what the authors call successful implementation. That 5%
is where the 95% comes from — it is the complement of one number in one category.

In the same chart, general-purpose tools like ChatGPT and Copilot run 80% investigated, 50% piloted,
and **40% successfully implemented**. Same report, same exhibit, eight times the success rate. The
headline picked the worse of the two and dropped the qualifier.

The methodology is 52 structured interviews, 153 survey responses collected at four industry
conferences, and a review of 300-plus publicly announced initiatives, between January and June 2025.
The report is candid about what that supports:

> These figures are directionally accurate based on individual interviews rather than official
> company reporting. Sample sizes vary by category.

That is an honest caveat, and it is doing a lot of work. "Directionally accurate, based on
interviews" is not a measurement. It is a well-informed impression, which is a legitimate thing to
publish and a poor thing to quote as a statistic.

Kevin Werbach at Wharton put the problem more bluntly than I would have:

> If MIT Project NANDA stands behind the claims, it should release the full supporting data. If not,
> it should retract the report.

<aside class="margin-note">
Kevin Werbach, Wharton, quoted in Futuriom, August 2025. The supporting data has not been released.
</aside>

**Verdict: contested, and rejected as usually stated.** There is no methodology behind "95% of AI
pilots fail". There is a 5% figure for one subcategory, sitting next to a 40% figure for another,
in a preliminary report that says its own numbers are directional.

## The second number

"Gartner predicts 30% of generative-AI projects will be abandoned after proof of concept by the end
of 2025."

This one is quoted as though the year has been graded. It has not. The statement is a **forecast,
published in July 2024**, about a year that had not happened yet.

<aside class="margin-note">
Gartner press release, 29 July 2024. Reached via the Internet Archive — the original returns 403 to
automated requests.
</aside>

Rita Sallam, the analyst quoted in it, was describing pressure rather than outcomes:

> After last year's hype, executives are impatient to see returns on GenAI investments, yet
> organizations are struggling to prove and realize value.

I could not find any retrospective — from Gartner or anyone else — establishing what the actual 2025
abandonment rate turned out to be. So the honest status is: a reasonable forecast, made by people
with good visibility, that nobody has checked.

Worth noting alongside it: Gartner's June 2025 prediction that over 40% of *agentic* AI projects
will be cancelled by end of 2027 names the same three causes as the 2024 one — escalating costs,
unclear business value, inadequate risk controls. Identical diagnostic categories, different
technology wave, a year apart. That tells you something about the categories.

**Verdict: plausible as a forecast, unverified as an outcome.** Cite it as a prediction or not at all.

## The third number, which is the instructive one

"47% of enterprise AI users have based a major business decision on hallucinated content." Usually
attributed to Deloitte. It travels with two companions: $67.4 billion in losses in 2024, and $14,200
per employee.

I tried to find it. The trail runs from one blog to another blog to a citation reading "Deloitte,
2025" — no report name, no page, no methodology.

So I went to Deloitte's actual *State of AI in the Enterprise* report, the 2026 edition, 3,235
senior leaders across 24 countries, the largest and most methodologically transparent survey in any
of this. **The figure is not in it.** Neither are its two companions.

<aside class="margin-note">
This is the most-cited number I could find with no traceable origin at all. The other two failed on
interpretation; this one failed on existence.
</aside>

**Verdict: rejected.** Not misinterpreted, not out of date — as far as I can tell it does not come
from anywhere. It looks Deloitte-sourced because a chain of posts said so, and each link in that
chain was written by someone who trusted the previous one.

## Two more, quickly, from the technical side

**"Lost in the middle."** The finding that a model degrades on information buried in the middle of
its context. It is the most-cited claim in context engineering, and it is real — for the models it
was measured on in 2023: GPT-3.5-Turbo, Claude 1.3, MPT-30B, LongChat-13B. All obsolete.

A replication published in May 2026 re-ran it on current open models and found the classic U-shaped
curve **mostly does not reproduce**; accuracy was, in their words, comparatively flat across
positions. So the field is citing a 2023 result about 2023 models as a description of how today's
models behave.

<aside class="margin-note">
Gabín, Pérez and Parapar, SIGIR 2026, testing Llama-3.1, Mistral-Nemo and Gemma-3. The original —
Liu et al., 2023 — remains the origin of the concept and good evidence about its own models.
</aside>

**"Context rot starts at 300–400K tokens."** The underlying research is real and useful: Chroma
tested 18 current models in July 2025 and found reliability degrading with input length. The
specific threshold numbers circulating in secondary coverage are not in their report. Someone
invented a round figure and it propagated.

Chroma also sells vector databases, and "long context is not a free lunch, you still need good
retrieval" is exactly what a retrieval vendor would want to be true. That does not make the data
wrong — it is multi-model and methodologically clear — but the funder belongs next to the finding,
every time.

## Why this keeps happening

None of these numbers spread because people are credulous. They spread because each one **confirms
something the reader already believes.**

Everyone in an enterprise has watched an AI initiative underdeliver. "95% fail" arrives as
vindication, and vindication does not get audited. The claims that survive unchecked are precisely
the ones nobody wants to check, which is the opposite of how it should work.

The tell is structural, and you can apply it without leaving your desk. A statistic that cannot name
its **population**, its **n**, and its **measurement** is not a statistic. All three of the failed
claims above collapse on the third one. Nobody measured whether an AI pilot "failed"; they asked
people at conferences how it was going.

## What actually survives

This is the part that matters, because the phenomenon is real even though the evidence for it is
not what people think.

**Adoption is broad and impact is shallow.** McKinsey's State of AI survey, published November 2025,
n=1,993 across 105 countries: 88% of organizations use AI in at least one function, while 39% can
point to any enterprise-level EBIT impact — and most of those put it below 5%. Self-reported, like
nearly all survey data here, but consistent across many independent reports.

<aside class="margin-note">
McKinsey, State of AI, 5 November 2025. I was unable to reach the primary page directly after four
attempts, so these figures rest on consistent secondary reporting rather than a fetch of the source.
Stated rather than hidden.
</aside>

**The model itself does work.** This is the finding that gets lost. Cui and colleagues ran genuine
randomized controlled trials at Microsoft, Accenture and a Fortune 100 company — 4,867 developers,
randomly assigned an AI coding assistant, **26% more tasks completed**. Measured output, not
self-report, published in Management Science.

<aside class="margin-note">
Cui, Demirer, Jaffe, Musolff, Peng and Salz, Management Science. The strongest methodology anywhere
in this piece.
</aside>

So the shape of the real problem is not "AI does not work." It is that a measured 26% individual
gain does not show up in the enterprise's accounts. Something between the individual and the P&L is
absorbing it.

The MIT report is actually good on this, in a section that got none of the attention its headline
did:

> The biggest thing holding back AI is model quality, legal, data, risk → What's really holding it
> back is that most AI tools don't learn and don't integrate well into workflows.

*Don't learn* is the phrase to sit with. A tool that cannot retain what it was told last week,
cannot be corrected once, and cannot carry an organization's own context into the next task will
produce exactly this pattern: real gains in the moment, nothing that compounds.

## The honest limit

I would like to tell you this proves the bottleneck is organizational rather than technical. It does
not, and the pieces claiming otherwise are overreaching.

Every source above is either a survey asking people to attribute blame — which is the kind of
question respondents are worst at answering — or a benchmark that says nothing about organizations,
or an RCT that measures individuals rather than balance sheets. **No controlled study isolates the
two variables.**

And there is real counter-evidence. METR tracks how long a task can be before frontier agents stop
completing it reliably. Agents succeed on tasks of a few minutes almost always, and on tasks beyond
roughly four hours less than 10% of the time. That horizon is doubling every seven months or so, but
it is doubling from a low base — and "complex, multi-week workflow" describes most of what an
enterprise actually wants automated.

So: capability limits are real at the long-horizon end. The organizational explanation is the
dominant reading of 2025–2026, and it is well-argued, but it is a reading rather than a
demonstration. I will take it as far as the evidence goes and no further.

## What to do with this

Ask three questions of the next AI statistic you are handed. Who was measured, how many of them, and
what instrument was used. If the answer to the third is "we asked them," you have an impression, and
impressions are fine as long as nobody calls them findings.

I keep a register of every number I use, with its source and its grade, because I have now been
wrong about one of them in public. That habit is not scrupulousness. It is the cheapest available
insurance against building an argument on a figure that turns out to be a chain of blog posts citing
each other.