FactoryXAutonomous AI WorkforceEnterprise Autonomy

Agent Washing on the Factory Floor

AM
Ajay Malik · Founder & CEO
August 28, 2026

Every vendor on the trade-show floor is selling an "agent" now. Most of them are selling the same rule engine and the same dashboard they sold five years ago, with a new word on the banner. There is a simple test that tells the two apart, and it has nothing to do with the demo.

A plant manager I know keeps a screenshot on her phone of a slide a vendor showed her last spring. The slide promised an "autonomous quality agent" that would "watch the line, understand deviations, and take action" across her injection-molding cells, and the demo was genuinely impressive — a clean interface, a natural-language summary of a scrap event, a confident little animation of the system "reasoning." She very nearly signed. What stopped her was a question she asked almost as an afterthought, standing in the parking lot afterward: when the agent decides a tolerance has drifted, what does it actually do next, on its own, before a human touches it? The answer, once she pushed past the marketing, was that it sent an email. It flagged the drift, wrote a nice paragraph about it, and routed that paragraph to a person, who then did every real thing that happened afterward. She had been shown a very expensive way to generate a better-worded alert, sold to her as a system that runs the line.

This is the texture of what is happening across manufacturing right now, and it is worth naming plainly because the confusion is costing real money. The floor is being flooded with software described as agentic — self-directed, autonomous, capable of acting — and a great deal of it is old technology wearing a new noun. The rule engines that have fired threshold alarms for twenty years are being recategorized as agents. The dashboards that have aggregated OEE and yield are being recategorized as agents. The industry has a word for the enthusiasm; it does not yet have a shared word for the disappointment that follows when the "agent" turns out to notice things and then wait, exactly as its predecessor did, for a human to do the work.

The industry already has a name for what most of this is

The tell is not subtle once you know to look for it, and the analysts have started saying so out loud. Gartner has predicted that over forty percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls, and among the specific culprits the firm names something it calls "agent washing" — the practice of rebranding existing tools, chatbots and robotic process automation and rule-based assistants, as agents without the underlying substance changing at all. That phrase deserves to travel further than it has, because it describes with precision the thing my acquaintance nearly bought. Agent washing is not a product category. It is a marketing operation performed on products that already existed, and manufacturing is an unusually easy target for it because the plant floor is already dense with systems that detect and report, which means the raw material for a convincing relabel is sitting there in every historian and every SCADA screen.

What makes the washing so hard to see through in a demo is that the front end has genuinely improved. Large language models have made it trivial to wrap a decades-old threshold rule in fluent, contextual language, so a deviation that used to surface as a red cell and an error code now surfaces as a paragraph that reads like a colleague explaining the problem. The intelligence appears to have moved into the system, when in fact only the narration has. Underneath the fluent summary, the logic is the same conditional it always was: if this value crosses that line, emit a signal. The system has learned to describe the problem beautifully and has learned nothing at all about resolving it. That gap between eloquence and agency is precisely where the money leaks, and it is invisible from the demo chair because a demo is a conversation, and these systems have become very good at conversation.

The test is whether it decides and acts, or merely detects and alerts

If you want to cut through the relabeling in a single question, ask what the system does after it knows something is wrong, before any human is involved. This is the whole distinction, and it maps cleanly onto a divide that has quietly governed plant economics for years: the gap between detection and action. A dashboard detects. A rule engine detects and then, at most, escalates. Both of them live entirely on the detection side of that gap, which is to say they deliver their value the instant a human happens to be looking and deliver nothing at all when no one is. A genuinely agentic system lives on the other side. It senses the deviation, reasons about what it means against the plant's own history and constraints, decides what should happen next, and then does it — pulling the maintenance record, checking the part against inventory, drafting the work order, proposing a maintenance window that respects the production schedule — pausing to put a decision in front of a person only when the decision genuinely belongs to one.

Those are not two points on a spectrum of the same capability; they are different capabilities, and only one of them has been automated in most of what is being sold as agentic AI in manufacturing. A rule engine that fires an alert is a dashboard with a louder voice, and a louder alarm is not autonomy no matter how gracefully it phrases itself. The reason the distinction matters so much on a factory floor specifically is that the expensive failures happen when no one is looking — at 3 a.m., between shifts, in the hours when a detected-and-alerted problem simply sits, correct and completely inert, waiting for a human to arrive and begin the response. A system that only detects has, by construction, done all of its work before the hard part starts. A system that acts is defined entirely by what it does in the hours the first kind of system spends waiting.

There is a second, subtler tell worth applying once a system passes the first, and it concerns how the system behaves when it encounters a failure nobody anticipated. A rule engine can only handle the deviations its designer wrote a branch for, which means the failure that scraps a shift — almost always the novel one, the combination nobody drew a rule around — falls straight through to a human every time, and the "agent" reveals itself as a lookup table with good manners. Real reasoning shows up exactly here, in the system's ability to take an unfamiliar signal, situate it against everything the plant knows, and compose a response that no one explicitly programmed. That is the difference between a fixed path and actual judgment, and it is the capability that the entire economic case rests on, because the anticipated failures were never the ones that hurt. The gains that make these systems worth buying — Deloitte has found that predictive maintenance can reduce unplanned downtime by 30 to 50 percent and cut maintenance costs by 10 to 25 percent — do not come from detecting faster. They come from closing the distance between knowing and doing, and a system that stops at knowing captures none of them.

What a plant that actually acts looks like from the floor

The practical way to hold all of this is to stop evaluating these systems by what they show you and start evaluating them by what they own. A washed agent owns the noticing; it will notice more, and describe what it noticed more articulately, and hand every consequence to you. A real one owns the outcome, which means it carries the work from the signal all the way to the resolution and only surfaces to a human at the points where a human's authority is actually required — the allocation call, the customer commitment, the decision that a machine should not make alone. This is the model behind platforms like StudioX's FactoryX, which runs specialist agents across the stages of production under a single reasoning core and an operating principle its designers state directly: you own the policy, the agents run the line. The human is not removed from the plant; the human is moved to the decisions that deserve a human, and everything mechanical between the decisions is absorbed. The test of such a system is not how well it talks in a demo. It is whether the line kept moving correctly at 3 a.m. while the policy you set was the only thing awake.

So the honest advice to anyone being shown an autonomous agent for their plant this quarter is to ignore the interface entirely and ask the parking-lot question my acquaintance eventually learned to lead with: after this thing understands that something is wrong, what does it do next without me? If the answer is that it tells you, more fluently than before, you have been shown a dashboard in a costume, and the label on the banner does not change what you will actually be buying. If the answer is that it decides and acts, under a policy you set and within gates you control, and comes to you only when the decision is genuinely yours, then you are looking at something new, and the distinction is not semantic. It is the whole difference between a plant that watches itself fail in well-written sentences and one that quietly keeps itself running while everyone who could have read the alert is asleep.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.