The pilot-to-production gap
The real, disclosed data on how many enterprise AI pilots fail to reach production — tiered honestly, including the study's own published critics
4 min read
This is the single most consequential statistic in enterprise AI right now, cited constantly and rarely tiered correctly. It traces to a real study with a disclosed methodology — which puts it well above the fabricated statistics this platform's other courses have had to name and discard — but the methodology has also drawn real, named criticism. Both things are true at once, and this lesson holds them together rather than picking the convenient one.
The study, named directly
In July 2025, MIT Media Lab's Project NANDA published "The GenAI Divide: State of AI in Business 2025," authored by Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari. Its headline finding, widely repeated across business and trade press: despite an estimated $30–40 billion in enterprise generative-AI spending, roughly 95% of organizations were seeing no measurable business return from their pilots — with only about 5% of integrated pilots extracting meaningful, rapid value. [Directional] — a real, named institution, a disclosed methodology, published in 2025; not [Established], for the reasons below.
The disclosed methodology: a systematic review of over 300 publicly disclosed AI initiatives, 52 structured interviews with organizational representatives, and survey responses from 153 senior leaders gathered across four industry conferences. [Established] as the study's own stated methodology.
The named criticism, held alongside the finding, not instead of it
The report drew real, specific pushback from credentialed reviewers, not just skeptics-in-general — and the honest thing to do with a study like this is name that criticism specifically rather than either repeating the 95% figure uncritically or dismissing the whole study because it's contested:
- Methodology-transparency criticism — at least one named academic reviewer (Kevin Werbach, a business-school professor who examined the report) publicly stated that if Project NANDA stands behind its headline claim, it should release the full underlying data, and that if it won't, the finding should be walked back. That specific ask — release the data, or retract — had not been fully answered as of this research.
- Sample-size and generalizability criticism — several reviewers pointed out that 52 interviews and 153 conference-attendee survey responses is a real but genuinely limited base from which to generalize a headline number about "enterprise AI" as a whole category.
- Potential-incentive criticism — because Project NANDA's own stated mission promotes agent-based, decentralized AI infrastructure as an alternative architecture, some reviewers noted the project has a plausible institutional incentive to characterize the current dominant enterprise-AI approach as failing.
- What survives the criticism, even from critics — multiple reviewers who challenged the precise 95% figure still endorsed the report's underlying diagnosis: that enterprise AI pilots fail primarily on integration, workflow fit, and the absence of a defined success metric before the build starts, not on model capability. That mechanism-level finding is considerably more durable than the headline percentage attached to it.
[Directional — contested] for the 95% figure itself; [Directional], better corroborated, for the underlying mechanism (integration and workflow failure, not model quality, is the dominant cause of stalled pilots), since that diagnosis is echoed independently across other enterprise-AI trade coverage, not just this one study.
What actually separates the roughly 5% that worked
The same report's own account of what distinguished successful deployments is worth taking seriously independent of the exact headline percentage, because it's a mechanism claim rather than a magnitude claim: pilots that succeeded tended to have a narrow, bounded scope, a clearly defined success metric set before the build started (rather than "let's try AI on this and see"), integration into an existing workflow rather than a stand-alone tool nobody's process actually routes through, and — echoing the previous lesson's build-vs-integrate point — externally sourced or vendor-partnered capability outperformed fully in-house builds, by the report's own account roughly two-to-one. [Directional — same study, same caveats above]
The honest synthesis for this course
Treat the specific "95%" figure the way this platform treats any single-study headline number: real, disclosed, worth citing by name with its source, and not something to repeat as settled fact without the caveat attached. Treat the underlying mechanism — narrow scope, a defined success metric before the build starts, real workflow integration, and a documented lean toward buying or partnering over fully in-house building — as the actually load-bearing, better-corroborated lesson here, and the one worth building a real go/no-go framework around (Module 5's kill-switch lesson uses exactly this shape).
Up next
Capital and team requirements
Said plainly: this is the most capital- and team-intensive business model on this platform, and the honest numbers behind that claim
3 min