AI MissionsAutonomous AI WorkersEnterprise AI StrategyupgradedEnterprise Autonomy

Decisions Per Person Per Week: A Better Metric Than Throughput

MW
Mark Weber · Chief Enterprise Architect
October 4, 2026

Every operations review measures how much got done. Almost none of them measure how much got decided — and in an organization where the routine volume increasingly runs itself, that is the number that actually tells you whether your people are doing anything worth their salaries.

In a Thursday operations review that could belong to almost any company, a director is presenting a slide with a number on it that has gone up. Tickets closed, invoices processed, cases resolved, units shipped, lines of code merged — the noun changes by department but the shape of the number does not, and the shape is always volume over time. The room nods at the number because it went up, and going up is understood to be good. Nobody in the room asks the question that would actually matter, which is how much of that throughput required a human being to think, and how much of it was a person moving work from one queue to the next, being the connective tissue between systems that refuse to talk to each other. The slide cannot answer that question because throughput was never designed to. It counts events, and it treats a judgment call and a copy-paste as the same event, worth the same tick on the same chart.

For most of the history of the modern enterprise, that was a forgivable simplification, because the two things were bundled together and could not be separated. A person who processed a hundred invoices a week was, somewhere inside that hundred, exercising judgment on the dozen that were ambiguous, and the throughput number was a rough proxy for the whole bundle. You could not buy the volume without also buying the judgment, so counting the volume was a reasonable stand-in for counting the value. That bundle is now coming apart, and when it comes apart the throughput number stops meaning what everyone in the room still thinks it means.

Throughput was a proxy that is quietly going stale

The reason to distrust throughput is not that it is wrong in some abstract sense. It is that it measures the wrong layer of the work, and it always did — it simply got away with it because nobody could act on the distinction. The data on how knowledge workers actually spend their time has been saying this for years, most legibly in software, where the activity is unusually easy to instrument. A SonarSource survey found that developers spend under a third of their time — around 32 percent — actually writing or improving code, and an IDC report put the figure as low as 16 percent, with the rest of the day going to coordination, maintenance, meetings, and the operational work of moving a change through an organization. If you were measuring an engineering team by commits and pull requests, you were measuring the smallest, and often the least deliberative, slice of what they did.

The same decomposition applies far outside software, because the pattern is not about code — it is about the difference between volume and judgment inside any role. A collections specialist, a claims adjuster, a leasing consultant, a support agent: run the same audit on any of them and you find a day that is mostly routing, reconciling, chasing, and remembering, with the genuinely difficult decisions occupying a surprisingly thin band of the hours. Throughput averages all of it together into one flattering number, and in doing so it hides exactly the thing a leader most needs to see, which is whether the expensive, hard-to-hire, easy-to-burn-out human in the seat is spending their week on the twelve ambiguous invoices or on the eighty-eight that a competent system could have closed without them. The number that goes up on the slide is silent on the only question that determines whether the headcount is well spent.

None of this mattered enough to fix while the volume and the judgment were welded together. It matters now because they are being pried apart, and the tool doing the prying is finally real enough that the routine layer can be lifted off the human entirely. When that happens, a team's throughput can stay flat, or even dip, while the value of what its people are doing rises sharply — because the volume they used to generate by hand is now generated without them, and the hours they have left are going almost entirely to the decisions. A metric that cannot tell those two situations apart is not a neutral simplification anymore. It is actively misleading, rewarding the teams that kept humans busy on routine and penalizing the ones that freed them for judgment.

What "decisions per person per week" actually counts

The alternative is not a slogan, and it is worth being precise about what it measures, because the precision is where its usefulness lives. A decision, in this frame, is a moment where a human applied judgment to something a system could not resolve on its own — an exception that did not fit the policy, a tradeoff between two commitments, a risk that needed a person to own it, a relationship call that no rule could make. Decisions per person per week counts how many of those moments each human handled, and deliberately does not count the routine volume that flowed around them. It is, in effect, the throughput of judgment, and it inverts the incentive of the old metric. Under throughput, a person looks most productive when they are personally touching the most work. Under decisions per person per week, a person looks most productive when the routine work has been lifted off them and what remains is dense with the calls that actually require a mind.

The reason this metric only becomes measurable now is that it depends on a workforce that can carry the routine layer autonomously, so that the decisions can be isolated and counted as a distinct thing. This is precisely what the move toward an autonomous enterprise makes possible: when Autonomous AI Workers read what comes in, gather the context from whichever systems hold it, and resolve the routine cases end to end — stopping to bring a human in only where the decision genuinely warrants one — the human's week reorganizes itself around exactly the set of moments the new metric is trying to count. In a platform built this way, with specialist agents running the volume under Human-in-the-Loop gates, every escalation to a person is, almost by definition, a decision, and the routine that used to inflate the throughput number never reaches a human to be counted at all. The architecture that makes the work autonomous is the same architecture that makes judgment measurable, because it is the thing that finally separates the two.

It is worth being honest that not every system sold as autonomous actually produces this separation, which is why the metric is also a useful lie detector. Gartner has predicted that over 40 percent of agentic AI projects will be canceled by the end of 2027, much of it what the firm calls "agent washing" — old workflows and chatbots relabeled without the underlying change in what they can do on their own. A relabeled workflow still hands every exception back to a person and still leaves them doing routine coordination between the steps, which means it moves the throughput number without moving decisions per person per week at all. If you are watching both numbers, the fraud reveals itself: real autonomy shows up as routine volume climbing while the human decision load stays lean and high-value, and the counterfeit shows up as a tool that made people faster at work they should no longer be doing.

Managing to the number that survives automation

Once you start watching decisions per person per week, a set of managerial habits that felt like common sense begin to look like relics of the bundled era. Staffing to volume — adding a head every time the workload grows — stops being automatic, because volume and human capacity are no longer the same variable; the question becomes whether the incremental work is decisions, which need people, or routine, which does not. Rewarding the busiest people stops being obviously right, because busyness under the old regime often meant a person had failed to offload routine they should have offloaded. And the health of a team stops being read off how much it processes and starts being read off how concentrated its human hours are in the moments that actually required a human — a team whose people spend their week almost entirely on judgment is not underworked because its throughput looks modest; it is operating exactly as an autonomous organization is supposed to.

The reframing to carry out of all this is that throughput was always a measurement of the old world, the world where a human's value and a human's volume were the same quantity because you could not purchase one without the other. That world is ending, unevenly and department by department, and as it ends the number on the Thursday slide is measuring less and less of what matters, quietly counting a machine's output and a person's judgment in the same undifferentiated tick. The organizations that see this will stop asking how much their people got through and start asking how much their people decided — because in a company where the routine runs itself, the throughput belongs to the software, and the only figure that still tells you anything about the humans is the density of judgment in their week.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.