Healthcare AIAI MissionsupgradedEnterprise Autonomy

An AI Mission for Healthcare: Prior Authorization

HE
Harry Edwards · Head of Solutions Engineering
May 21, 2025

Prior authorisation is filed under paperwork, which is why it keeps producing clinical problems. Almost none of the waiting is a decision being made — and the part that isn't a decision is the part software should be carrying.

It is twenty to seven in a specialty clinic and the last patient left forty minutes ago. Someone on the clinical staff has a payer portal open in one window and two years of chart notes open in another, scrolling backwards through encounters looking for the visit where the earlier treatment was started, the visit where it was noted to be failing, and the outside record that arrived as a scanned document and was never indexed in a way that makes it findable now. The therapy itself was decided that morning, in the room, by a clinician with the patient in front of them. What is happening at twenty to seven is not that decision being revisited; it is the slow reconstruction of evidence that the decision was reasonable, assembled into the sequence and the vocabulary that a set of published criteria happens to ask for. Nothing in this hour is difficult. All of it is unavoidable, all of it is being done by someone whose training was for something else entirely, and until it is finished the request does not exist as far as the payer is concerned.

That last point is the one that gets lost when prior authorisation is discussed as an administrative burden. The clock that matters to the patient does not start when a reviewer opens the file; it starts when the clinician decides the patient needs the treatment. Everything between those two moments — the gathering, the formatting, the resubmission after a portal rejects a packet for a missing field, the phone call to an outside practice for a record that should have arrived with the referral — is time the patient spends without care, and it is time that almost nobody is measuring, because it accrues in the space between two organisations where neither one's reporting reaches. A delayed authorisation is not an inconvenience filed under overhead. It is a clinical event with a clinical cost, borne entirely by the person who is not in the room when it happens.

Almost none of the wait is deliberation

Watch enough of these requests move and the shape becomes hard to unsee. The elapsed time from decision to determination divides into two very unequal parts, and the smaller part is the one everyone argues about. Adjudication — a qualified clinical reviewer reading a complete file and forming a medical-necessity judgement — is genuine expert work, and when the file is complete it does not take long. Assembly is everything else: locating the prior therapy history, pulling the imaging report and the relevant labs, finding the note that documents why a first-line option was not appropriate for this particular patient, translating all of it into the structure the criteria expect, submitting it, and then discovering what was missing and doing a second lap. The overwhelming share of the wait is the second category, and the second category contains no judgement at all.

It is worth being precise about what "no judgement" means here, because the phrase does a lot of work. It does not mean the assembly is unimportant — a packet missing the one note that establishes the clinical picture will produce a wrong answer, and a wrong answer here has consequences. It means the assembly is determinate. The criteria state what evidence is required. The record either contains that evidence or it does not. Deciding whether a document is the prior therapy record the criteria are asking for is a retrieval and matching problem against a body of enterprise knowledge; it is not a judgement about whether the treatment is warranted, and the person currently doing it is not making that judgement either. They are searching. They are searching in a system that was designed to record care rather than to prove it, in fifteen-minute fragments between patients, or after hours, on top of a day that was already full.

This is why the burden question is not evenly distributed. When a request is delayed or comes back adverse, the work of fixing it falls to the two parties least equipped to absorb it: the patient, who is unwell by definition and who is being asked to navigate an appeals process while unwell, and the clinical staff, whose time is finite and whose next patient is already waiting. Neither of them created the gap. Both of them pay for it, and the patients who fare worst are consistently the ones with the least capacity to advocate — those without someone to make the calls for them, those whose records are scattered across practices that do not share systems, those for whom another two weeks is not a scheduling annoyance but a deterioration.

Assembly can be automated; the answer cannot

Here is where the discussion about AI in this workflow usually goes wrong, and the error is not subtle. Because assembly and adjudication sit inside the same process and produce the same artifact, it is tempting to treat them as one automatable pipeline — to reason that if a system can read the criteria well enough to gather the evidence, it can read them well enough to apply them. That inference is false, and the harm that has already been documented from automated adverse determinations is the reason to state it flatly rather than hedge it. A system may gather, structure, verify completeness, and present. A qualified clinical reviewer decides. There is no version of this where a model issues a denial, an adverse determination, or a medical-necessity judgement, and no efficiency argument that makes that acceptable.

The asymmetry between the two error modes is what makes the line hold. If an assembly system mis-files a document or fails to find a record, a human sees an incomplete packet and goes looking — the failure is visible, local, and correctable before anything reaches a patient. If an adjudication system is wrong, a patient does not get treatment, and the failure is invisible precisely to the person harmed by it, who receives a determination that looks identical whether it was reasoned or merely generated. One failure produces rework. The other produces a clinical outcome. Building for the first while refusing the second is not caution for its own sake; it is the only design in which the speed gained is unambiguously good for the patient, because every minute removed is a minute of retrieval, not a minute of thought.

There is a market reason for the discipline as well as an ethical one. Gartner has predicted that more than forty percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear value, and inadequate risk controls. In most domains an over-scoped deployment fails commercially and quietly. In this one it fails on a person. The systems that will still be running in five years are the ones whose scope was drawn where the accountability naturally sits, which in utilisation review means everything up to the determination and nothing past it.

A queue where every remaining minute is deliberation

What that scope looks like when it is built well is less dramatic than the pitch usually suggests and considerably more useful. Autonomous AI workers read the criteria set the request will actually be measured against and work backwards from it, searching the record the way the person at twenty to seven was searching it but across every source at once — encounters, outside documents, imaging, labs, the medication history — assembling a candidate packet with each element traced back to the document it came from. Where evidence appears to be absent, the system says so explicitly rather than papering over the gap, because an assembly layer that guesses is worse than none. Where a document is ambiguous, it surfaces the ambiguity to a person rather than resolving it. This is the human-in-the-loop pattern that the wider move toward autonomous enterprise operations has been converging on across industries, and platforms in this space — StudioX among them — tend to describe the boundary in similar terms: specialist agents run the gathering, humans own every decision with a consequence attached. In this workflow the consequence has a face, so the boundary is not a configuration option. A qualified clinical reviewer decides, on a complete file, with the reasoning attributable to a person.

The payoff is not a faster denial, which would be a worse outcome dressed as an improvement. It is that the request arrives complete on the first submission, that the second lap disappears, that the clinician's evening is returned to them, and that the reviewer spends their attention on the question they are qualified to answer rather than on chasing a missing record. It also produces something the current process almost never yields: a legible account of what was known, what was requested, and when, which is exactly the material an appeal needs and exactly what patients are currently expected to reconstruct themselves.

So the mental model worth carrying is a stopwatch with two hands. One measures the time a request spends being deliberated on by a person qualified to deliberate; the other measures the time it spends waiting to be assembled. Today the second hand runs for days and the first for minutes, and the entire public conversation is about the first. The goal is not to make deliberation faster — deliberation is the part that deserves the time, and compressing it is how patients get hurt. The goal is a queue in which the only remaining minutes are deliberative ones, where the waiting that is left is a person thinking carefully about a complete file, and where nobody is spending an evening proving something they already knew to be true.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.