Most Enterprise AI Pilots Are Quietly Dying
How MIT's research on 300 AI deployments found a 95% failure rate, and what separates the rare pilots that actually pay off.
Opinion. Most companies running generative AI pilots this year will have nothing to show their board for it, and the reason has almost nothing to do with which model they picked.
MIT's Project NANDA, a research initiative at the university's Media Lab, published a study in 2025 titled The GenAI Divide: State of AI in Business, based on 150 executive interviews, a survey of 350 employees, and an analysis of 300 public AI deployments. Its central finding, still widely cited through 2026 as enterprises plan next year's AI budgets, is stark: roughly 95% of generative AI pilots deliver no measurable return on investment, according to reporting on the study by Fortune, via Yahoo Finance. Only about 5% of pilots achieve rapid revenue acceleration.
Aditya Challapally, the report's lead author, told Fortune that the divide is not primarily a model-quality problem. "The 95% failure rate for enterprise AI solutions represents the clearest manifestation of the GenAI Divide," the report states, and the actual cause is what Challapally calls a "learning gap," in both the tools companies deploy and the organizations deploying them. Consumer tools like ChatGPT succeed for individuals precisely because of their flexibility. That same flexibility becomes a liability inside an enterprise, where a tool that doesn't learn a company's specific workflows, data, and edge cases stalls out well short of the process it was meant to automate.
The two decisions that actually predict success
The MIT research isolates two variables that matter more than model choice, and both are uncomfortable for engineering-led AI strategies. First, buy versus build: companies that purchased AI tools from specialized vendors and partnered on integration succeeded roughly 67% of the time, while companies that built their own AI systems internally succeeded at about a third of that rate. That gap holds even in financial services and other regulated industries where in-house development is often treated as the safer, more controllable default.
Second, where the budget goes. More than half of surveyed generative AI budgets are allocated to sales and marketing tools, the highest-visibility, most board-friendly use cases. But MIT found the largest actual returns in unglamorous back-office automation: eliminating outsourced business processes, cutting external agency spend, and streamlining internal operations. The pilots getting funded and the pilots generating returns are, for most companies, two different lists.
Why this keeps happening despite the warning
The GenAI Divide report is not new information at this point, it has been widely cited for over a year, and enterprises are still repeating the pattern it describes. That persistence is itself informative. Executives are not failing to notice the failure rate; they are structurally incentivized to fund visible pilots over effective ones. A customer-facing chatbot pilot is easy to demo to a board. A back-office workflow automation project that quietly eliminates a business-process-outsourcing contract is not a good slide, even if it is the one that pays for itself.
There is also an organizational tell buried in the data: companies were, in Challapally's account, reluctant to share their own failure rates even with MIT's researchers. That reluctance suggests the 95% number is not a secret failing to spread through the industry. It is a number companies would rather not confront internally, because confronting it means admitting that the pilot the CEO announced at the last all-hands is one of the 95%, not the 5%.
Stakeholder takeaways
CFOs and finance leaders evaluating AI budget requests for next year should treat "which model are we using" as close to the least important question in the room. The buy-versus-build ratio and the specific workflow being targeted predict outcomes far better than model selection, and both are decisions made before a single line of code gets written.
Line managers, not central AI labs, are the actual lever MIT's research points to for successful adoption. Pilots owned and driven by the managers closest to a specific workflow outperform top-down initiatives run entirely out of a centralized innovation team, because line managers are the ones who know where the actual friction is.
AI vendors selling point solutions have a real data-backed argument to make against internal build-it-yourself efforts, but only for the narrow, well-scoped, back-office automation category where MIT found returns concentrated. The same vendors overselling flashy customer-facing use cases are pitching into the part of the market most likely to fail, a dynamic that echoes broader shifts in how enterprise AI budgets actually get spent.
Workers in customer support and administrative roles are already experiencing the disruption MIT documented, largely not through mass layoffs but through unfilled vacancies as positions previously earmarked for business-process outsourcing quietly disappear, a pattern consistent with what Edgewisely found in hospital AI pilots that also struggle to pay off.
The bigger picture
The AI industry has spent two years selling transformation and delivering, by MIT's own measurement, a 5% hit rate. That is not evidence AI doesn't work. It is evidence that most companies are running the experiment backwards, funding the pilot that looks good in a keynote instead of the one that fixes a real, expensive, boring problem. For teams choosing which agent framework or platform to build on, this data is a reminder that picking the right production tooling matters less than picking the right problem. The 5% of companies getting this right are not smarter about AI. They are more honest about what actually needs fixing, and willing to buy a narrow tool that fixes it instead of building a broad one that doesn't.
Frequently Asked Questions
What did MIT's GenAI Divide report find?
MIT's Project NANDA found that roughly 95% of generative AI pilots at enterprises deliver no measurable return on investment, based on 150 executive interviews, a survey of 350 employees, and analysis of 300 public AI deployments, while only about 5% of pilots achieved rapid revenue acceleration.
Why do most enterprise AI pilots fail?
According to the report's lead author, Aditya Challapally, failure is driven mainly by a "learning gap," not model quality: tools and organizations that don't adapt to specific company workflows stall out, and budget is often misallocated toward visible use cases like sales and marketing tools rather than higher-return back-office automation.
Is it better to buy AI tools or build them internally?
MIT's data shows companies that purchased AI tools from specialized vendors and partnered on integration succeeded about 67% of the time, roughly three times the success rate of companies that built AI systems entirely in-house.
Where do enterprise AI pilots actually generate the most value?
The report found the biggest returns in back-office automation, such as eliminating outsourced business processes and cutting external agency costs, rather than in the customer-facing and sales and marketing tools that receive the majority of AI budgets.
Editor's note — sources: Sheryl Estrada, "MIT report: 95% of generative AI pilots at companies are failing," Fortune, via Yahoo Finance, August 18, 2025 (updated coverage cited through 2026); MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025."