AI GovernanceRegulated IndustriesComplianceupgradedEnterprise Autonomy

AI Governance for Regulated Industries

AM
Ajay Malik · Founder & CEO
August 12, 2025

Regulated industries already know how to prove that a decision was defensible. The recurring mistake in AI governance is treating that standard as something new to invent, rather than something already written and waiting to be applied to an actor who happens not to be a person.

An examination usually begins in an unremarkable room with an unremarkable request. The examiner has pulled a sample — forty decisions from a single quarter, chosen without much ceremony — and works through them one at a time, asking the same short list of questions about each. Who made this decision. What information were they looking at when they made it. Under what authority were they permitted to make it at all. And, the question that quietly governs the other three, can the file prove it. For thirty-eight of the forty, this goes fine, because the institution has spent decades building machinery whose entire purpose is to answer exactly those questions about a human being: the sign-off, the second reviewer, the contemporaneous note, the delegation matrix that says which title may approve which amount, the retention schedule that keeps all of it findable years after everyone involved has moved on. Then the examiner reaches the two decisions that were made by the system, and the conversation changes tone. Not because anyone believes those decisions were wrong. Because what the institution can produce for them is a log, and a log is not a file.

That gap — between a log and a file — is where most AI governance programs quietly fail, and it is worth being precise about why, because the failure is almost never a failure of intent. The institution took the risk seriously. It stood up a committee, adopted a set of principles, built an inventory of models, commissioned a fairness assessment, and wrote a policy that reads well. What it did not do was ask, before anything shipped, what the record of a single automated decision would need to look like when a stranger examined it eighteen months later without access to anyone who was there. It treated governance as a discipline to be assembled alongside the system, when the standard it was actually being held to had been sitting in the institution's own procedures the whole time.

The examiner's questions did not change; only the actor did

The four questions are old, and they are not really about technology at all. They are the operational form of a much simpler principle: that a consequential decision affecting someone must be attributable, explicable, authorized, and provable after the fact by evidence created at the time. Underwriting works this way. So does dispositioning a complaint, approving a trade above a limit, adjusting a claim, signing off a batch release, or making a call that touches a patient's care. Every one of the controls that surround those acts exists to convert a human judgment — which is private, fallible, and rapidly forgotten — into something a third party can reconstruct later without relying on anybody's memory. The industry settled this argument a long time ago: recollection is not evidence, the file is.

Once you see it in those terms, the peculiar thing about most AI governance work is how much of it is spent reasoning about the system in the abstract rather than about its acts in particular. Model documentation describes what a system is designed to do, over a population, at a point in time. A fairness assessment describes aggregate behaviour across cohorts. A risk taxonomy places the system in a tier. All of this is useful and none of it answers the question actually being asked in the room, which is why this claim, belonging to this person, was denied on this date, on what information, and by whose authority. The examiner is not auditing a category of system. They are auditing a decision, and a decision is a singular event with a specific factual basis that either was or was not captured while it happened. You cannot retrofit contemporaneity. Whatever the system failed to record in the moment is gone in the same way a reviewer's unwritten reasoning is gone, and for the same reason.

This mismatch also explains the organizational pattern that regulated firms keep reinventing and keep being disappointed by. When governance is constituted as a parallel discipline — its own function, its own vocabulary, its own committee sitting outside the delivery path — it necessarily arrives at the end, and functions that arrive at the end produce documentation rather than evidence. They can approve or refuse; they cannot reach back into a runtime that has already been built and cause it to have captured something it was never designed to capture. That is why so many of these programs end up as a queue of assessments blocking deployment while the underlying evidentiary problem goes untouched, and it is part of why the analyst view of the current wave is so unsentimental. Gartner has predicted that more than forty percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. In a regulated setting, "inadequate risk controls" rarely means nobody thought about risk. It means the thinking never turned into a record anyone can stand behind.

Designing the file is designing the system

The productive inversion is to treat the audit file as a specification rather than an output — to write, before a line of the system is committed, the document you would hand an examiner for one automated decision, and then build whatever is required to make that document true. It is a mundane exercise and it is remarkably clarifying, because each field in it turns out to be an engineering requirement in disguise.

Take the basis of the decision. The file must show the information the system actually had at the moment it acted, which is a stricter thing than the information it could in principle have retrieved. A system that consults enterprise knowledge dynamically produces a different factual basis on Tuesday than on Thursday, so the evidence has to include what was retrieved, from which source, in which state, alongside the version of policy in force at that instant. Take authority, which is the field most deployments have no answer for at all. Human actors derive their authority from a role and a written delegation with limits attached, and nobody in a regulated firm finds this exotic; yet very few AI deployments are accompanied by an equivalent instrument saying that this system may act unsupervised within these categories, up to these thresholds, and must escalate beyond them, granted by a named accountable person on a specific date. Without that, the honest answer to "under what authority" is that the system acted because somebody deployed it, which is not an authority, it is an accident of scope.

Then there is the human checkpoint, which deserves more skepticism than it usually gets. A recorded approval is only evidence of review if the reviewer could plausibly have exercised judgment — if they saw the material facts, had the time, and had a real option to refuse. A queue that asks one person to clear several hundred items a shift, each presented as a summary with an approve button, generates a pristine audit trail attesting to a review that did not meaningfully occur, and that is a worse position to be in than having no checkpoint at all, because the record now asserts something the institution cannot defend. Human-in-the-loop is a design commitment about what the human is shown and how much of it they can absorb, not a checkbox in an architecture diagram.

What all of this implies is that the record cannot be a byproduct scraped from application logs after the fact. It has to be a first-class output of the runtime, emitted with the same reliability as the action itself, which is a platform property rather than a policy one. It is the reason the more serious enterprise systems in this category are converging on a similar shape — autonomous AI workers that execute missions inside the enterprise's own deployment boundary, retaining their observations as the durable, inspectable trace of what was perceived and concluded at each step; a single gateway through which every model call passes so that attribution is structural rather than reconstructed; human gates positioned at the points where authority genuinely runs out. StudioX is built on that premise, and the reason is not primarily a compliance one: a system whose reasoning can be reconstructed by a stranger is also a system its own operators can debug, tune, and trust. The body of work now accumulating around the autonomous enterprise keeps arriving at the same place from different directions — that evidence produced at the moment of the act is what separates a deployment that survives contact with an examiner from one that gets quietly shelved.

None of this is legal counsel and none of it substitutes for the specific obligations any given institution carries; regimes differ, supervisors differ, and the mapping from these fields to those obligations is work that belongs to the people who own it. But the design instinct travels everywhere, and it is worth stating as plainly as possible. A system is governed when any one of its decisions can be reconstructed by someone who was not there, without asking anyone who was — the same test the institution already applies to its people, applied without special pleading to a participant that happens to be software. Judged by that test, the useful question is not whether an AI governance framework exists. It is whether you can write the examination transcript in advance. If you can draft the file for a decision the system has not made yet, you have governed it. If all you can draft are the principles, you have described how you would like the system to behave, and left the proving to a room you have not yet been invited into.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.