IA 360
Current Affairs

The 95% AI-pilot claim: what the NANDA report measured

NANDA’s preliminary report says 95% of enterprise generative-AI solutions show no financial impact. It publishes its scope and limitations, but not the denominator needed to reconstruct that figure.

6 min read AI-generated Leer en español
The 95% AI-pilot claim: what the NANDA report measured

On 18 August 2025, a striking number began to circulate: 95% of enterprise generative-AI pilots deliver no return. It came from The GenAI Divide: State of AI in Business 2025, a report of preliminary findings produced in collaboration with MIT-linked Project NANDA. The document says 95% of organisations are getting zero return and only 5% of integrated pilots are extracting millions of dollars in value.

Yet the report does not publish the exact numerator and denominator that produce that percentage. Nor does it consistently use one unit: it refers to organisations, pilots, solutions and tools. The honest conclusion is not that the figure is false. It is that this is the authors’ estimate, and readers cannot reconstruct it from the displayed data. Learning to detect that difference—between a stated figure and a reproducible rate—is the capability that survives after the headline expires.

What NANDA actually studied

The authors—Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari—describe a multi-method design. From January to June 2025, they reviewed more than 300 publicly disclosed AI initiatives and announcements, conducted structured interviews with representatives of 52 organisations and collected responses from 153 senior leaders at four industry conferences. The document describes itself as preliminary findings, not a randomised trial or a census of every business.

The appendix adds an important time window. It defines success as deployment beyond the pilot stage with measurable indicators, and says ROI was measured six months after the pilot and adjusted for department size. It also says confidence intervals were calculated through bootstrap resampling “where applicable.” However, it provides no table connecting each initiative to its status, indicators, observation window and outcome. It does not display a confidence interval for the 95% either.

The body offers another definition. For task-specific generative tools, successful implementation means one that users or executives reported as causing a “marked and sustained” impact on productivity or profit and loss. The first definition requires deployment and indicators; the second incorporates interviewees’ assessments. They may overlap, but they are not interchangeable.

The missing denominator

The adoption chart says that, among embedded or task-specific generative-AI tools, 60% were investigated, 20% reached a pilot and 5% were successfully implemented. The following text calls this the “95% failure rate for enterprise AI solutions.” In the executive summary, however, the 95% refers to organisations with zero return, while the 5% refers to integrated pilots extracting value.

We do not know whether the denominator is the 300-plus public initiatives, the 52 interviewed organisations, a subset of evaluated tools or a combination. The counts behind each bar are not published. It is therefore inaccurate to restate the finding as “95% of all AI pilots fail,” or to say the 5% accelerates revenue. The report discusses productivity or P&L impact, deployment and value—related but distinct concepts.

This rewrite also removes the estimate of $30–40 billion in investment that opens the document. The report states it but does not identify a source there or explain how it was calculated. A striking figure does not become traceable merely because it appears beside an academic logo.

What the observations do show

The qualitative pattern is useful. Interviewees described rigid tools with insufficient contextual memory, weak integration into daily workflows and a continuing need for manual intervention. The report calls this distance the learning gap. A model answering well in isolation is not enough: an enterprise system must retain permitted context, integrate with data and permissions, receive corrections and work repeatedly.

This explains why individual productivity and financial return are not synonyms. Saving ten minutes on an email may be real, but it does not automatically cut costs or increase revenue. To reach P&L, an organisation must multiply the saving by frequency and adoption, subtract review, errors, licences, integration, support and risk, and demonstrate that freed capacity found an economic use. Without that bridge, a task improvement is not yet an enterprise return.

The document also found a difference between implementation routes. In its interview sample, external partnerships using customised tools reached deployment about 67% of the time, compared with 33% for internal builds. But the limitation appears beside the chart: definitions varied, the observation period could be short, precise initiative volumes were unavailable, and organisations that buy may differ in capabilities, risk tolerance or procurement. This is correlation in a 52-organisation sample, not proof that buying causes success.

The limitations are not small print

The methodological appendix acknowledges that the sample may not represent all regions or enterprise segments; organisations willing to talk may differ from non-participants; and success metrics vary across businesses and industries. It adds that public projects may miss private developments, simultaneous operational and economic changes complicate attribution, and six months may understate results from complex enterprise systems.

Those warnings change the reading. The 300-plus public initiatives help reveal known narratives and deployments, but an announcement does not necessarily contain full costs or audited results. Interviews add context, although cases and quotations are anonymised for confidentiality. The survey describes people who responded at those conferences, not a random sample of the world’s businesses.

The institutional relationship also deserves exact wording. The report says it was produced “in collaboration with Project NANDA” and that its views belong solely to its authors and reviewers, not affiliated employers. The current MIT Media Lab NANDA page describes the group and its work on networked agents; the former official report URL now redirects there. The copy linked in this article is an archived capture of the PDF that was publicly hosted on NANDA’s domain.

How to measure a pilot without manufacturing a percentage

Before starting, a company should fix six elements. First, the unit: a task, team, process or profit-and-loss account. Second, a baseline from before the system. Third, a comparable group or period to estimate what would have happened without the tool. Fourth, a window long enough for the process. Fifth, every cost: licence, integration, data, review, incidents and training. Sixth, quality and risk limits that prevent purchasing speed with errors.

It can then write the equation. Net benefit is verifiable additional revenue plus avoided costs, minus the system’s total cost. ROI divides that net benefit by total cost. If the organisation measures hours, it must declare how saved time becomes money: capacity actually reassigned, external spending avoided or additional volume—not salary multiplied automatically.

The results sheet should preserve counts, not percentages alone: how many pilots started, how many reached the deadline, how many withdrew, which threshold each was meant to cross and how many crossed it. It should separate “produced no return” from “could not yet be measured” and “cancelled for another reason.” Only then does a 5% or 95% have a denominator another reader can audit.

A useful number as a question, not a verdict

The NANDA report is right to identify a common gap between a demonstration and an integrated system. It also provides its own caveats, which many headlines lost. What it cannot support is turning 95% into a general law about all enterprise AI.

Whenever the next failure rate travels, ask for five items: observed unit, numerator, denominator, definition of success and time window. If one is missing, the estimate can be reported with attribution and limits, but not converted into a universal fact. Here the faithful wording is straightforward: the authors of a preliminary NANDA report estimated a 95% rate; the document presents valuable signals about integration, but does not publish the counts needed to reproduce it.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close