GovernmentAI MissionsupgradedEnterprise Autonomy

An AI Mission for Government: Defensible Case Review at Scale

PG
Patrick Gilberg · Head of Accounts
April 24, 2025

A public office can decide a case well and still be unable to show that it did. The distance between those two things is where public trust is actually lost, and closing it has almost nothing to do with deciding faster.

The file is four years old and it takes a while to locate, because the archive is organised the way archives are organised — by identifier, by category, by the year the case was opened. When it finally comes up, it contains everything a file is designed to contain: the forms as they were submitted, the correspondence in the order it arrived, the supporting documents, the determination itself, the date it was issued, and the initials of the office that issued it. What the file does not contain is the reasoning. Nothing in it records which of the two or three plausible readings of the circumstances the office settled on, or which parts of a thick and partly contradictory submission actually carried weight, or what the office had done in the small number of earlier cases that looked very much like this one. The person who worked it has moved to another part of the service. Someone reads the file three times and comes away able to say precisely what happened and entirely unable to say why, which is the only question the review was convened to answer.

None of this implies the original decision was wrong. It very probably was not; most of them are not, and the people who make them are careful in a way that rarely gets acknowledged. But a decision whose soundness cannot be demonstrated is, in administrative terms, hard to distinguish from a decision that was never reasoned at all, and the citizen on the other side of it has no way to tell those apart either. That is the peculiar exposure of public administration. The work is done under conditions of enormous care and recorded under conditions of enormous haste, and it is the record, not the care, that has to survive contact with a reviewer years later.

What a reviewer is looking for is not correctness

It is worth being exact about what "defensible" means in this setting, because the word gets used loosely and its actual meaning is narrower and stranger than it sounds. A defensible decision is not one that has been shown to be correct. Very often correctness is not even available as a standard, because the circumstances were genuinely ambiguous and two conscientious people reading the same file could reasonably land in different places. What a reviewer can assess is something else: whether the reasoning that produced the decision can be reconstructed from the record, and whether that reasoning was applied consistently to cases that resemble this one. Those two properties — reconstructable and consistent — are what "defensible" actually names, and a decision can possess both while remaining, in some ultimate sense, arguable.

This is not a legal technicality dressed up as a principle. It is the operating condition of any administration that treats people as equals, because equal treatment is not a claim about outcomes in isolation but a claim about the relationship between outcomes. A person who is told no has been treated fairly if others in materially the same position were told no for the same articulated reasons, and has been treated unfairly if they were not, even where the individual determination looks impeccable in isolation. That relational quality is invisible inside any single file. It only becomes visible when the file can be set beside its neighbours, and setting a file beside its neighbours is precisely the thing that is nearly impossible to do in most public offices today.

Consistency is asserted far more often than it is checked

Ask how an office knows that like cases are being treated alike, and the honest answer in most places is that it knows through people. Long-serving staff carry the precedents in their heads, remember the awkward case from a few years back that established how a particular pattern of circumstances gets handled, and pass that memory along informally to newer colleagues over months. This is a real and underrated form of institutional knowledge, and it is also undocumented, unevenly distributed, and mortal: it leaves when people leave, it does not transfer to a new region or a new team, and it cannot be audited, because the precedents were never written anywhere a reviewer could find them.

The record systems do not fill the gap, because they were built to answer a different question. They can tell you every case of a given category opened in a given period, which is what a management report needs, and they cannot tell you which prior cases involved a similar configuration of circumstances, which is what consistency requires. To find the like cases you have to already know they exist. So consistency ends up being asserted in good faith rather than demonstrated, and when it fails it fails quietly, in ones and twos, distributed across regions and years in a pattern nobody has the means to see.

The failures do not land evenly, and this is where the stakes stop being administrative and start being personal. A decision that goes wrongly against a person is not the mirror image of one that goes wrongly in their favour; the first can cost someone housing, income, status, or time they will not get back, while the second is a cost the state absorbs and can usually recover. The formal answer to the first is appeal, and appeal is a genuine safeguard, but it is an unequal one, because mounting one demands exactly the reserves — time, money, literacy in the process, health, the confidence to push back against an institution — that the people most likely to be wrongly refused are least likely to have. Which means the correction mechanism cannot be relied on to surface systemic inconsistency. The record has to do it, and at present the record mostly cannot.

The thing to automate is the record, not the determination

Here is the line that should not move. No software should make, recommend, or effectively pre-empt a determination that changes a person's entitlement, status, or legal position. A human caseworker makes that decision, signs it, and owns it, including its consequences for the person it lands on. This is not caution for its own sake, and not a transitional arrangement to be relaxed once the technology matures, because accountability is only meaningful if it attaches to someone who can be asked to explain themselves — and a system that produces an outcome for a person to ratify has moved the decision while leaving the accountability behind.

But the scarce commodity in that four-year-old file was never the decision. It was the record — and building a record is exactly the kind of patient, unglamorous, high-volume work that a caseworker does badly at the end of a long day and that machines do tirelessly. Two functions are worth separating. The first is contemporaneous assembly: capturing what was actually in front of the caseworker at the moment of decision, which documents and which versions, which prior contacts, what was requested and what was missing, so the reconstruction years later is a matter of reading rather than inference. The second is retrieval before the fact rather than after it: surfacing the earlier cases whose circumstances resemble this one, with what was decided and the stated reasoning, at the point where the caseworker is still deciding rather than at the point where an auditor is questioning.

That second function is the one that changes the character of the work, and it has to be built with some discipline to avoid becoming the thing it must not be. A system that ranks the comparable cases by how strongly they point toward a particular outcome has begun making the decision through the side door, because presenting a conclusion is the most reliable way to produce agreement with it. Retrieval that supports rather than substitutes shows the neighbouring cases and their reasoning without proposing an answer, and asks the caseworker for one thing the current process almost never asks for: where this case departs from its neighbours, say so in writing. Departure is entirely legitimate — circumstances differ, and rigid consistency would be its own injustice. Unexplained departure is the problem, and the fix is a sentence typed by the person who decided, not a score generated by a model.

This is the distinction that the body of work published on enterprise autonomy keeps returning to: autonomous systems earn their place in high-stakes settings by owning the work around the judgment rather than the judgment itself. It is also why the discipline matters more here than almost anywhere. Gartner has predicted that over forty percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear value, and inadequate risk controls — and inadequate risk controls in a commercial setting produce a written-off investment, while in a casework setting they produce citizens who were treated differently from their neighbours and cannot find out. Framed as an AI Mission on a platform like StudioX, the shape is deliberately modest: specialist agents assemble the file and retrieve the comparable history from enterprise knowledge, Human-in-the-Loop is not a checkbox at the end but the determination itself, and every observation the agents made is retained so that the reasoning available to the caseworker on the day is still available to a reviewer in a decade.

The reframing worth carrying out of this is that the case file has been misclassified for as long as public administration has existed. It has been treated as residue — the paperwork left over once the real work of deciding is done — when it is in fact the only durable product the office makes. The determination is consumed by the person it affects on the day it arrives; the record is what the institution can still stand behind long after everyone involved has moved on. An office that understood this would stop measuring itself only by how many cases it closed and start asking a harder question about the ones it has already closed: how many of them could be explained, today, to someone who was not there, without having to find the person who signed them. In most places that number would be uncomfortable. It is also, unlike caseload and cycle time, a number that can be moved without asking anyone to decide anything faster.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.