AI Automation
spotting agent washing
On the difference between AI bolted on and agent-native, and the questions no vendor can answer with a slide
Thousands of vendors call themselves agentic AI. Thousands. And how many of them actually deliver? Around 130, Gartner reckons. The rest are polished-up chatbots and repackaged process automation with a new word on the box. Gartner stuck a name on it back in June 2025, agent washing, and by now the term just sits there in the legal reference works. In the average vendor demo, though? You’ll never hear it. And here’s the thing. If you’re buying software today you can’t sit around waiting for a regulator to do the spadework for you. There isn’t one here yet. So you set the bar yourself. A handful of questions no vendor wriggles out of with a pretty slide.
Over in the 95 percent myth I hold the failure numbers around AI projects up against the method underneath them. This one? This is the buying side of that same story. Simple as this: how do you keep your project from starting on a promise the product is never going to keep?
what exactly is agent washing?
You stick a new label on software you already had, agentic AI, when there’s nothing materially agentic about it. That’s the whole trick. Gartner describes exactly that, vendors giving AI assistants, chatbots and RPA tooling a fresh coat of paint. And in that same press release they predict over 40 percent of agentic AI projects die before the end of 2027. Why? Rising costs, unclear value, risks nobody’s really keeping a grip on (Gartner, 2025).
That 130 figure, out of the thousands who call themselves this? Same release. And I went looking for the method behind it. Came back with nothing. No denominator, no measurement setup, nothing. Shame, because that’s exactly the bit that would make the number worth something. So treat it as a ballpark from an analyst firm. It isn’t a real measurement. And the secondary sources, well, they cheerfully turn it into “95 percent is fake”. That percentage is nowhere in the primary text, mind.
And why is any of this even possible? Simple reason. There’s no fixed definition to hold anything up against. Look, even MIT Sloan says it straight out, a broadly shared definition of agentic AI just isn’t there (MIT Sloan, 2025). The rule of thumb you hear everywhere, generative AI answers a prompt and an agent perceives, reasons and acts on its own toward a goal, that one’s fine. Only every vendor draws its own little maturity ladder right next to it. And an empty box where the definition should be? That’s an open port for marketing.
how do you tell bolted-on ai from agent-native?
I’ll admit it straight off, I’ve got an unfair advantage here. I build both kinds, every single day. Rule-based automation in Make scenarios and monday workflows, for the CRM landscape of an installation firm I run. And real agents over MCP connections, on that same landscape, plus fikst, my own agent-native CRM. So the difference between a workflow somebody calls an agent and an agent that actually does multi-step work on its own, I know that one from the inside.
Look, a Make workflow that can’t independently call different tools with reasoning power, that isn’t an agentic workflow. That’s more of an automation, or just an integration hookup between two systems: A happens and it does B. You might have ten variants of A and ten variants of B. But actually calling different tools on its own, with reasoning, a thing like that doesn’t do it.
So that’s my own line for it. And it drops you straight at the first test, take the AI out and look at what’s left. Two tests, really, and for both of them you honestly don’t have to be technical at all.
The removal test first. Say the AI feature vanishes from the product tomorrow. Does anything essential change about what the system can do? No? Then that AI was a chat layer sitting on top of an otherwise unchanged package. Ten variants of A, ten of B, same as it ever was. That’s it.
The second one I call the equivalence test. Simple question: can that agent do the same as a human behind the screen? Actually create, change, schedule. And at best set up the system itself, instead of just summarizing and dropping a suggestion in your lap. HFS Research pours that same check into a slightly more formal two-gate test, agency and scalability, and warns very specifically about copilots that get relabeled as agents while all they really do is text-to-action inside workflows you already had (HFS Research, 2025).
And why does that matter so much? AI bolted on means every new automation is custom work all over again. From scratch. The chat layer can chat about your data, fine. Act inside your process? Can’t do it. Agent-native means the actions themselves are within reach for the agent, and then that second and third automation suddenly isn’t a fresh project anymore. You never spot that difference in a demo, by the way. Makes sense, a demo always shows you the prettiest side. You spot it in the architecture. And so, in the answers to your questions.
which questions do you ask a vendor?
Six of them. You don’t need more. And you just ask them in plain words, right across the table. The answers tell you more than any product video ever will.
- Can the thing actually do the action, or does it just advise me? Ask it dead concrete. Can it create an order, move an appointment, add a field? And in which systems has it genuinely got write access, right now?
- Show me a live demo on my own case. Not a recorded clip, not some rehearsed script. My customer, my process, right here on the spot.
- Where do I see back what the agent actually did? A serious product logs it, every action. HFS says it flat out to buyers: make them show the evidence. Test results, audit logs, run traces, don’t just swallow the claims. And honestly? This is the one I’d fight hardest for of the six.
- What happens when it gets it wrong? Can I undo the action, are there permission boundaries, and when does it hand things back to a human?
- How do you measure the success rate? Watch this one. fin.ai, a player in this market themselves, mind you, flags two tricks: blended figures where human and machine get counted as one score, and customers who give up without ever being helped but still land in the “resolved by AI” pile (fin.ai, 2026). So keep pushing for the number where no human stepped in.
- Does the platform even tell human users and agents apart? Separate permissions for agents, that’s a sign the thing was actually built for them. And not bolted on as an afterthought later (MindStudio, 2026, vendor-affiliated, but concrete enough to put to them).
A vendor who rattles through these without breaking a sweat doesn’t need to slap a single label on anything, far as I’m concerned. And the one who reaches back for that slide with the word agentic on it? They’ve answered you too. Just not the way they were hoping.
where does it go wrong when nobody presses?
And when nobody presses? Eventually it lands in court. That’s not a thought experiment, mind. Over in the United States the files are already stacking up. Since September 2024 the FTC’s been running cases against misleading AI claims under the banner Operation AI Comply, everything from the “robot lawyer” DoNotPay to tools that sat there writing fake reviews (FTC, 2024). The SEC fined advisers who bragged about AI they didn’t even use, and in January 2025 it took on its first listed company, Presto Automation: that “proprietary” speech tech turned out to belong to a third party, and most orders still ran through human hands (SEC, 2025). And the founder of shopping app Nate? Criminally prosecuted since April 2025. The app promised neural networks. The orders got done by hand, by contract workers (Holland & Knight, 2025). Line those cases up and you’re basically looking at the mirror image of the six questions above: technology claims that didn’t hold, automation numbers blown out of proportion, human hands quietly kept off-camera.
And two of the famous horror stories deserve the exact same skepticism you’d aim at a vendor, by the way. Take the viral one, Builder.ai “had 700 engineers pretending to be AI”. Not true. I went and read the most careful reconstruction of it, and it shows something else entirely: a small AI team, a thick layer of perfectly real outsourced development that got sold as AI speed, and a bankruptcy that in the end came out of revenue fraud. Not the AI fiction itself (Pragmatic Engineer, 2025). And Amazon’s checkout-free stores? Also not just “a thousand people instead of AI”. There genuinely was a working model, and the fight was about the hidden scale of the human review (The Verge, 2024). Look, the lesson lands a touch sharper than the meme does. A human in the loop is dead normal in serious AI. The red flag is the vendor who won’t tell you honestly how that ratio really sits.
and in the netherlands?
Not one case on agentic claims yet. Far as I could dig up, anyway. Closest thing is a ruling from the Advertising Code Committee about AI-generated songs sold as “personally composed” (Stichting Reclame Code, 2024). Real, but narrow. Dutch AI oversight is carved up across ten-odd bodies, with the ACM holding the consumer-deception piece, and the transparency duties from the AI Act don’t kick in until 2 August 2026 anyway (ICTRecht, 2026). Can the ACM handle this? On greenwashing it’s already shown it’ll come down hard on vague claims, and that playbook drops onto AI claims pretty neatly in theory. But today that’s still an analogy. Not practice.
So act as if a regulator’s never turning up. Those six questions cost you half an hour in a sales meeting. A project that starts on a promise the product can’t keep? That costs you a year. And to be clear, I build agent-native software myself. Go on, hold my work to that exact same bar. That’s what it’s for.
frequently asked
- What is agent washing?
- The term was coined in June 2025 by Gartner: vendors repackaging existing chatbots, AI assistants or RPA software as agentic AI while the product cannot carry out multi-step actions on its own. Gartner estimates that of the thousands of self-declared agentic vendors, only around 130 actually deliver substantial agentic capabilities; no exact methodology behind that figure has been published, so read it as an order of magnitude.
- How do I check whether software is really agentic?
- Two tests work in any sales conversation. The removal test: mentally take the AI feature out of the product; if everything still works exactly the same, it was a chat layer, not an agent. And the equivalence test: can the agent perform the same actions as a human through the screen, including creating, changing and scheduling, or can it only advise? Beyond that, always ask for a live demo on your own scenario and for the log of what the agent has done.
- Is a human in the loop a sign of fake AI?
- No. Nearly every serious AI system uses people for training, review and escalation; with Amazon's checkout-free stores the debate ultimately turned on the concealed scale of that human role, not on its existence. The red flag is not that people are watching, but a vendor unwilling to state the ratio between human and machine honestly.
- Will a regulator step in on misleading AI claims?
- In the United States it has for years: the FTC has been running enforcement cases under Operation AI Comply since September 2024, the SEC fined Presto Automation among others, and the founder of shopping app Nate is being criminally prosecuted because the 'AI orders' turned out to be processed by hand. In the Netherlands there is no case on agentic claims yet; AI supervision here only becomes active from August 2026. Until then nobody filters for you and the test sits with the buyer.