An AI Mission for Expense Management

Finance departments review nearly every expense claim that arrives, and nearly every claim is fine. The interesting question is not how to check them faster — it is which ones were ever worth checking.
A finance analyst spends a good part of a Friday afternoon inside an expense system, opening reports one after another. There is a taxi fare from an airport to a hotel with a photographed receipt attached, correct to the cent. There is a lunch for two with the attendees named, comfortably inside policy. There is a train ticket, a parking charge, a bag of conference-booth supplies bought from a hardware store because the shipment did not arrive. Each takes somewhere between forty seconds and two minutes to open, read, compare against a policy document running to a few dozen pages, and approve, and none of them will turn out to be wrong. Somewhere in the same queue, buried among hundreds of identical small approvals, sits one claim that genuinely merits a second look — a duplicate submitted twice under slightly different descriptions, or a category that has been quietly drifting above its budget for two quarters. The analyst will probably find it, and they will find it having paid for the privilege with several hours of attention spent on things that were never in question.
Almost every finance organization runs this way, and almost nobody has done the arithmetic on it out loud. The unspoken premise of expense control is that reviewing everything is prudent — that the process pays for itself in what it catches and deters. But a control is not automatically worth its cost simply because it is a control, and expense review is unusual in that its cost is easy to underestimate and its yield easy to overstate. The cost is diffuse: a few minutes from an analyst, a few from a manager signing off between meetings, a few more from the employee who fills in a field, gets a query, and answers it. The yield is concentrated and memorable: the one report that was wrong, retold at the next process review as evidence the whole apparatus is working.
The control costs more than the exceptions are worth
Consider the shape of an expense population rather than its individual members. Claims are distributed the way most operational things are distributed — a very long tail of small, ordinary, entirely compliant items, and a small head of larger, more complicated ones where the money and the ambiguity both live. Review effort, by contrast, is almost flat. It costs roughly the same amount of human attention to open, read, and approve a modest taxi fare as it does a mid-sized travel report, because most of the effort is in the opening, the context-loading, and the closing rather than in the judgment itself. Put a flat cost against a long tail and the outcome is arithmetically inevitable: the overwhelming majority of the organization's review labour is spent on the portion of spend where almost none of the risk is, and it is spent there permanently, every cycle, forever.
The comparison that ought to be made and rarely is: what does review of the tail actually recover? Not what the whole control recovers — the head of the distribution, the genuinely complex claims, plausibly justifies careful attention many times over. The tail is a different question. If a category of claim is small enough that even a total loss on it is immaterial, and reliable enough that findings are vanishingly rare, then reviewing that category is a net transfer of value out of the business. It consumes analyst hours that could go to forecasting, vendor negotiation, or accrual quality, and manager hours — among the most expensive a company buys — on a task requiring no managerial skill whatsoever. It also consumes something less visible but more corrosive: the credibility of finance with everyone else, because an employee queried about a nine-euro coffee learns something about where the function's attention goes, and it is not a lesson that makes them keener to engage.
None of this is an argument that expense policy does not matter or that controls are theatre. Policy matters enormously, and the deterrent effect of a credible review process is real. But the deterrent effect is a function of credibility, not coverage — of the well-founded belief that something unusual will be noticed, which a good risk-weighted process often delivers better than universal review does. Reviewing everything at a uniform, shallow depth is a poor deterrent precisely because depth is what catches anything; a reviewer with two minutes per report and four hundred reports is not really looking at any of them. The organization ends up paying for comprehensiveness and getting the detection power of a glance.
Universal review is an inherited habit, not a considered design
It is worth asking where the everything-gets-checked default came from, because it was not chosen on the evidence. It came from the era when approval was a physical signature on a paper form, when there was no way to route a claim conditionally because there was no routing at all — there was an in-tray, and the manager worked through it. Universal review was not a risk posture; it was the only thing the medium could express. Every subsequent generation of expense software reproduced that shape in digital form, adding receipt capture, card feeds, and mobile submission, while leaving the fundamental question — who looks at this, and why this one — exactly where the paper in-tray left it.
Finance already knows better, incidentally, and applies better logic elsewhere without noticing the inconsistency. No competent audit function tests every transaction; it stratifies a population, samples the low-risk strata, and concentrates effort where materiality and risk indicate. Nobody regards this as negligence, because it is not — it is the correct allocation of a scarce resource against a known distribution. Expense management is the one place where an organization full of people who understand sampling and risk-based testing reviews the whole population at uniform depth, and it does so mostly because that is what the workflow does and nobody has proposed anything else.
The proposals that do get made almost always take the wrong form. Faced with a review burden, the natural instinct is to make review faster: better receipt scanning, tighter policy rules in the submission form, a dashboard showing the approval backlog, a rule that auto-approves anything under a fixed amount. The last of these is closest to the right idea and still misses it, because a flat threshold is not risk-weighting — it is a blunt guess that treats a small claim from a pattern-free submitter identically to a small claim that is the fourth in a suspiciously regular series. Everything else on that list optimizes the wrong verb, because making a review that should not be happening take ninety seconds instead of two minutes does not recover the review; it just makes the waste more efficient.
Weighting attention is a different machine than accelerating review
What risk-weighted attention actually requires is not speed but judgment applied to every claim before anyone decides whether a human should see it — a genuinely different capability from anything a rules engine offers. It means reading the claim in context: what this person's spending looks like over time, what their team's pattern looks like, whether the trip it belongs to has other claims attached, whether the amount is unremarkable in isolation but part of a shape that is not, whether the policy language covering it is clear or has a known grey edge. A fixed rule cannot hold that much context, which is why rules-based expense systems end up either so permissive they catch nothing or so strict they generate a flood of false queries that trains everyone to click through them. Deciding what deserves attention is reasoning work, and until recently there was nothing available to do it except the reviewer whose time you were trying to protect.
This is where a good deal of current enthusiasm goes wrong, and it goes wrong in a predictable way. Gartner has predicted that over forty percent of agentic AI projects will be canceled by the end of 2027, citing unclear business value and what it calls "agent washing" — familiar tooling relabeled without any change in what it can actually do unattended. An expense product that extracts a receipt total more accurately, or drafts a politer query email, is exactly that: assistance layered onto a process whose fundamental design flaw is that the process exists for the whole population. The change worth making is structural. A system built as an ongoing mission rather than a form-filling aid takes standing responsibility for the claim population — forming its own observations about spending patterns over time, pricing the attention each claim deserves against the risk it actually carries, clearing the compliant long tail on its own authority, and escalating to a human the specific cases where judgment is genuinely required, with the reasoning attached so the human starts from context rather than from a blank report. That posture, in which software owns the execution and people own the decisions and the policy, is the substance of what the category publication Enterprise Autonomy documents as the shift from assisted work to autonomous work, and it is the operating model behind platforms like StudioX, where a mission runs continuously against a stream of work and human-in-the-loop approval is reserved for the points where a person's judgment changes the outcome.
The mental shift this asks for is small to state and awkward to accept, because it means saying out loud that some spending will not be looked at by a person and that this is the correct outcome rather than a tolerated lapse. Stop thinking of expense review as a checkpoint that every claim must pass and start thinking of it as an attention budget that has to be allocated. The budget is real and finite — it is the hours of analysts and managers, and every hour spent confirming that a compliant taxi fare was compliant is an hour not spent on the claim, the category, or the vendor relationship that would actually have repaid the scrutiny. Organizations that make this shift are not loosening control; they are concentrating it, moving from a thin uniform film of oversight spread across everything to real depth where depth changes something. The question stops being how quickly finance can get through the queue and becomes how well the queue was chosen — and once it is put that way, the current queue, which contains nearly everything, is very obviously not an answer anyone would arrive at by design.
Discussion
No comments yet — start the conversation.