Back to home

AI Automation

the 95 percent myth

On the most quoted AI statistic of the moment, and what remains of the failure numbers once you hold every percentage against its method

I build and run AI automation for smaller companies. A monday CRM landscape for an installation firm, and fikst, my own agent-native CRM where real agents work on live business software through MCP. That crossing, from a demo that shows well to something that runs in production every single day with all the management hanging off it, I made that myself. So every time another slide floats past with the headline that 95 percent of AI projects fail, I want to know where that number got measured. Comes from an MIT report. And the methodology? Told inconsistently even in the coverage that’s selling it. The other failure numbers doing the rounds? Same kind of holes. A small sample, data nobody’s allowed to see, or a vendor footing the bill. And still. Even the solid sources point the same way. A small minority of companies demonstrably makes money with AI. The pattern holds. The percentages wobble.

In the production gap I lay out the thesis I’m checking here against the numbers: the demo is 20 percent of the work, and most projects still die somewhere between pilot and production. So the question here’s simple. Where do these failure numbers come from, and how much is left once you hold each one up against its own method?

is it true that 95 percent of ai projects fail?

Nobody knows. And the MIT report the number comes from? It sure doesn’t prove it. August 2025, one headline went round the world: 95 percent of generative AI pilots deliver no measurable effect on the profit and loss statement, and that despite 30 to 40 billion dollars of investment (Fortune, 2025). That headline came out of “The GenAI Divide: State of AI in Business 2025” by MIT Project NANDA. Nearly a year on and it’s still in just about every strategy deck I watch go by.

Want to read the report yourself? Fill in a request form first. The raw data was never released. And the methodology gets told all over the place. The report itself talks about 52 structured interviews, 153 surveys collected at four conferences, and a review of more than 300 public AI initiatives. Fortune’s own coverage says 150 interviews and 350 surveys (Futuriom, 2025). Same coverage, same report, and the counts just don’t match each other. Sounds like a detail, right? It isn’t. Analyst site Futuriom went further, laid the chart the figure leans on next to the conclusion, and asks whether that 95 percent can even be pulled out of it. Wharton professor Kevin Werbach put careful, public question marks next to the report too. And the harder demand, that MIT release the full data or pull the report? That one’s Futuriom’s own. Two really different things, mind you. In the retellings they nearly always get squeezed into one.

Maybe the real number’s lower. Maybe higher. The most quoted AI figure of the moment is also the least checkable one, and that’s the honest answer. A report about sloppy AI rollouts that itself looks methodologically sloppy. That irony? Worth saying out loud.

where do the other failure numbers come from?

From research that’s almost always smaller, softer or more self-interested than the headlines let on. Walk down the best-known ones and you keep hitting the same gap. The percentage in the headline, and the measurement sitting under it.

Take the second most popular one, from RAND Corporation. More than 80 percent of AI projects supposedly fail, twice as often as IT projects without AI (RAND, 2025). Read the publication itself and it literally says “by some estimates”. Look, RAND is leaning there on somebody else’s guess. Its own research was something else entirely: interviews with 65 experienced data scientists and ML engineers, and out of that a qualitative analysis of five root causes. Valuable work. Just doesn’t back up that percentage. And in just about every retelling the two get glued together like they were one measurement.

Then there’s the doubling that did the rounds in 2025. 42 percent of companies scrapped the majority of their AI initiatives that year, against 17 percent in 2024, and on average 46 percent of proofs of concept died before production. Those come from the “Voice of the Enterprise” series by S&P Global Market Intelligence, a survey of more than a thousand IT and business decision makers across North America and Europe. The series is real, and several independent sources cite the exact same numbers. Getting into the primary publication myself, though? Didn’t manage it either. So I pass them on the way I found them, confirmed secondhand (WorkOS, 2025).

In June 2025 Gartner made a prediction. More than 40 percent of agentic AI projects get canceled before the end of 2027: rising costs, unclear business value, weak risk controls (Gartner, 2025). A prediction, so by definition you can’t test it against the outcome yet. But there’s a side catch in that same press release I honestly find more interesting. Of the thousands of vendors calling themselves agentic AI, only about 130 actually deliver agentic capabilities, Gartner reckons. The rest sell polished-up chatbots and RPA. Gartner christened that “agent washing”. And that word on its own is worth holding onto at your next vendor demo.

And the Dutch number: 21 percent AI project success, lowest of six European countries studied, and at the same time the highest internal resistance (38 percent). That’s from “State of Integration & AI 2026”, research commissioned by integration platform Frends and run by Sapio Research (Frends, 2026). Two things to sit with. An integration platform has a real stake in an integration-first approach coming out on top. And the sample only starts at organizations of 201 employees and up. Big companies, then, and the top end of the midmarket. For the installer with fifteen on the payroll? This number says almost nothing.

what survives once you strike the shaky numbers?

One pattern. And methodologically independent sources confirm it over and over: pretty much everyone’s got AI in use, and the measurable gains stay with a minority. That one rests on sturdier stuff than the headline percentages above.

Morgan Stanley tracks every quarter how many S&P 500 companies can name a measurable AI benefit. 10 percent end of 2024, 15 percent in the third quarter of 2025, 21 percent end of 2025 (Morgan Stanley, 2026). The line’s climbing. And it’s still a minority, at the biggest and best-funded companies on the planet no less. IBM put the question to two thousand CEOs early in 2025: a quarter of the initiatives delivered the expected ROI, and 16 percent got scaled company-wide (IBM via Fortune, 2025). McKinsey, consultancy, so keep that in mind, lands on the same picture: 88 percent of organizations use AI in at least one business function, 39 percent see measurable impact on the bottom line, and roughly 38 percent are past the pilot phase (McKinsey via CX Today, 2026). And BCG, also advisory, counts a quarter of 1,800 executives reporting significant value (BCG, 2025). Advisory work too, sure.

For small business the OECD survey’s the most relevant, more than two thousand SMBs across twelve countries. AI adoption among SMBs grew from 7.1 percent in 2023 to 17.4 percent in 2025. And at the same time the gap with the big companies actually got wider, from 23.4 to 34.6 percentage points (OECD, 2026). Eurostat lands on 55 percent AI use at large companies in the EU, against 17 percent at the small ones (Eurostat via Ipsos, 2026). Those I keep EU-wide on purpose, mind. A reliable Dutch SMB figure I haven’t found, and what’s floating around on Dutch marketing blogs contradicts itself so hard that I use none of it.

Add it up. Equity analysts, a CEO survey, two consultancies, official statistics. All different methods, and they all land in the same spot. So you don’t need to believe MIT’s 95 percent at all to take the production gap seriously. If anything: whoever builds their whole story on that one contested percentage is doing exactly what the report accuses companies of. Impressive demo, thin backing.

what do the companies that do capture value do differently?

Same handful of moves keep coming back in the research. Each with its own source, and its own caveat.

First one, buying over building. That same contested MIT report has a finding that got way less attention than the headline: bought-in AI solutions and partnerships came off in roughly 67 percent of cases, in-house builds in roughly 33 percent (Fortune, 2025). Same report, so same caveat. But I recognize it straight away. I build fikst myself and run a bought-in platform alongside it, and it’s exactly the self-build side where management, maintenance and the road to production get underestimated. Every time.

Second, starting small and staying focused. That quarter of companies pulling significant value at BCG bets on a handful of initiatives, scales them fast, and reworks the underlying core processes while it’s at it (BCG, 2025). Consultancy research again, with a consulting practice as its stake. And it rhymes with the IBM picture: the scaling is the bottleneck.

Third, data and integration in order before the pilot. Network vendor Cisco calls 13 percent of companies “Pacesetters” in its AI Readiness Index, and that group turned four times as many pilots into production (Cisco, 2025). Vendor research, same as that Frends report spotting the same pattern across Europe. More independent is the study of agentic AI at industrial companies that Forbes cited this month, with a term that nails exactly what I hit once agents have to run on real business systems: the “capability-deployment verification gap”. The pilot does fine in the test setup. And the moment it hits live business data? The confidence is gone. Just like that (Forbes, 2026).

how to handle failure numbers and pilots yourself

Say you’re a small-business owner and an AI proposal lands on your desk tomorrow. This is the order I keep myself.

First thing, I distrust every round percentage. I ask about the sample, about who paid for it, and whether the raw data’s public. A number out of vendor research? I treat it as marketing until proven otherwise. Then I pick one process where the euros or hours are already measurable: quotes, customer questions, planning, invoicing. And before that pilot starts I pin down what success is, in money or hours a week, and who’s going to measure it. Only then does the build question come up. My preference is to buy or rent a proven solution before you have anything built, because the best-backed success patterns point that way. But the step I weigh heaviest myself sits further down the row, and it’s the one that gets skipped most: check whether the data the system needs actually exists, is right, and is reachable. Small as a tidy FAQ, big as your ERP integration. And plan the road to production before the pilot kicks off. So: who manages it later, what’s it cost per month, what happens when it breaks. If it stalls anyway, stop in time. A three-week pilot you call off costs a fraction of a year-long zombie project.

What can you safely skip? Training your own model. Standing up a company-wide AI transformation program. Signing with every outfit that’s slapped the word agent on its homepage. Not one number in this essay gives you a reason to. Gartner’s agent-washing estimate points the other way, if anything.

where I stand

I don’t believe the 95 percent. And the reverse hype? I don’t buy that one either. Pumping failure numbers around is its own kind of demo culture: an impressive headline sat on a thin measurement, while the real pattern underneath needs no headline at all. Independent sources keep showing the same thing. Measurable AI value stays with a minority, and that minority buys in, starts small, and has its data in order before the pilot starts. Duller than a number like 95, sure. But it’s the one of the two you can actually build a company on.

frequently asked

Is the MIT study saying 95% of AI pilots fail reliable?
It is the most quoted and at the same time most contested AI figure of the moment. The raw data was never released, the sample is reported inconsistently (52 or 150 interviews, 153 or 350 surveys) and critics such as Futuriom dispute whether the percentage follows from the published charts. Use it at most as an illustration of a broader pattern, never as hard fact.
So how many AI projects really fail?
A reliable overall percentage does not exist. Independent sources do converge on the same picture: 21% of S&P 500 companies could name a measurable AI benefit at the end of 2025 (Morgan Stanley), 25% of AI initiatives delivered the expected ROI according to 2,000 CEOs (IBM) and 39% of organizations see measurable impact on the bottom line (McKinsey). Measurable value stays with a minority, even without an exact failure rate.
Why do AI pilots fail so often?
The recurring causes in the research: no success metric agreed up front, data that is not in order, building in-house where buying would have done, and pilots never designed to reach production. For agentic AI, Gartner adds "agent washing": vendors selling ordinary chatbots as agents, so projects start on a promise the product cannot keep.
What can I do as a small business to stay out of the failed projects?
Start with one process where the hours or euros are measurable, agree up front what success means and buy a proven solution instead of building your own. Check your data before the pilot starts and decide immediately who will manage the system afterward. And dare to stop the moment the agreed metric is not met.
nlen