An AI Mission for Banking: AML Alert Triage
Financial crime teams have spent a decade trying to get better at deciding. The decision was never the expensive part — the assembly of evidence behind it was, and that is the part a machine can finally carry.
An analyst sits down on a Tuesday morning with a queue of transaction monitoring alerts and already knows, in the way that experience makes a thing knowable before it is proven, how most of the day will end. The overwhelming majority of what is in front of her will close as no further action. Not because she will be careless, and not because the scenarios that generated the alerts were badly built, but because monitoring rules are calibrated to be generous by design: they are supposed to catch things, and catching things at scale means catching a great many ordinary customers doing ordinary things that happen to resemble something else. So she opens the first alert, and then she does what will occupy nearly all of her working day — not judgment, but retrieval. She pulls the customer's profile from one system and the account's history from another. She reads back through prior alerts on the same party to see whether this pattern has been looked at before and what was concluded. She checks the counterparties, the stated occupation, the expected activity captured at onboarding, the source-of-funds notes some colleague wrote eighteen months ago in a free-text field. She screens the names. She reconstructs a timeline. And then, having spent forty minutes assembling a picture, she looks at it for perhaps ninety seconds, decides it is consistent with what the bank already understood about this customer, and writes three sentences explaining why.
Then she does it again. And again, several hundred times a month, across a team, across a year. That is what AML alert triage actually is in most institutions, and describing it as an analytical function badly misrepresents where the hours go. It is an evidence-assembly function with a short analytical step at the end.
The cost of the negative finding is what actually drains the function
Financial crime programs measure themselves almost entirely on the positive: the suspicious activity identified, the case escalated, the report filed, the account exited. Those are the outputs that justify the function's existence, and it is right that they get the attention. But the economics of the function are not set by the positives. They are set by the negatives — by the vast body of alerts that will close cleanly and produce, as their only artifact, a documented explanation of why nothing happened. Every one of those costs the same evidentiary work as a case that goes somewhere. The analyst cannot know in advance which alert is which, so she builds the full packet for all of them, and the file that ends in no further action is, in labor terms, nearly indistinguishable from the file that ends in an escalation.
This is why staffing an AML function feels like pouring water into sand. Volume grows with the customer base, with every new product, every new corridor, every tuning change that widens a threshold because someone found a gap. Each increment of volume converts directly into hours, because the unit of work is not a decision — decisions are fast — but a reconstruction, and reconstructions are slow no matter how skilled the person doing them. Hiring more analysts buys more reconstructions. It does not change the ratio of assembly to judgment, which is the actual pathology. Neither does better tuning, which is worth doing on its own merits but only ever moves the volume; the packet behind each surviving alert still has to be built by hand.
And the work itself is corrosive in a specific way that anyone who has managed one of these teams will recognize. You hire people for their instincts about financial crime, tell them their judgment is the reason they are there, and then hand them a job in which judgment occupies a small fraction of their attention while the rest goes to opening tabs, copying identifiers, and paraphrasing the same facts into the same narrative structure for the thousandth time. The good ones get bored, and bored is dangerous in a control function, because the failure mode of boredom is not obvious error. It is the quiet drift toward pattern-matching the shape of an alert rather than reading it, which is precisely the condition under which the one that mattered slides past.
What is automatable is the packet, not the call
Once you separate the two halves of the task, the automation question stops being philosophical and becomes almost mechanical. Assembling the evidentiary packet is retrieval, correlation, and drafting across systems that already hold every fact required. The customer record, the account activity, the prior alert history and its dispositions, the screening results, the onboarding profile, the relationship notes, the public-record context — none of that requires a person to exist. It requires something that can reach into each of those systems, understand what is relevant to the specific scenario that fired, resolve the entities so that the same party in four systems is recognized as one party, order the events into a coherent timeline, and write the first draft of a narrative that explains what the activity appears to be and how it does or does not fit what the institution already knows about this customer. That is a substantial piece of cognitive work, but it is not a decision. It produces a document, and the document is the same document a careful analyst would have produced, arrived at faster and, importantly, arrived at consistently, without the variance that creeps in when the same task is performed by twelve people on the third hour of a Friday.
Most of what is currently marketed into this space does not do that. It scores the alert, or it clusters alerts, or it presents a dashboard in which the analyst can see more things at once while still doing all the assembly herself. This is the gap that makes so much of the current wave disappointing to the people who buy it — Gartner has predicted that over forty percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear value, and what it calls "agent washing," the rebadging of older tools as autonomous without the underlying capability changing. A model that ranks an alert has not removed a single minute of packet assembly. It has only told the analyst which reconstruction to perform first.
Structured as an AI Mission, the shape is different. Specialist agents work the retrieval across core banking, the monitoring platform, case management, screening, and the institution's own Enterprise Knowledge — the policy language, the scenario documentation, the historical dispositions that encode how this bank has actually reasoned about this pattern before. A Reasoning Core assembles what they return into a single coherent packet with its observations traceable to their sources, drafts the narrative in the institution's own voice and format, and flags the specific tensions a human should look at: the fact that contradicts the customer profile, the counterparty that appeared in a prior alert, the gap in documentation. What arrives at the analyst is not an alert. It is a finished file with a recommendation and a visible chain of evidence, and her ninety seconds of judgment now sits at the front of the task instead of the end of it. This is the same reallocation that runs through the broader shift toward autonomous enterprise operations in every function where skilled people spend their days being connective tissue between systems that will not talk to each other.
The disposition has to stay with a person, and that is a feature
It would be tempting, having automated the packet, to let the same system close the alert. Resist that, and not only for the reasons you would expect. Supervisors of financial institutions generally expect that a suspicious-activity decision is owned by an accountable individual within the institution — that the file shows who looked, what they considered, and why they concluded what they did, and that the person named there could be asked about it years later and answer. A control function is not merely a process that produces correct outcomes; it is a process that produces a record of responsibility. An automated disposition, however well reasoned, produces an outcome without an owner, and an outcome without an owner is not a control at all, which is why Human-in-the-Loop here is not a transitional safeguard to be engineered away in a later release. It is the point at which the whole architecture is anchored.
That constraint turns out to be clarifying rather than limiting. It tells you exactly where the boundary belongs: everything up to and including the drafted recommendation is machine work, and the act of disposition — the moment a person reads, agrees or disagrees, and signs — is human work, permanently. It also tells you what the system's real deliverable is. The deliverable is not a decision. It is a decision-ready file, and the measure of whether the technology is working is not how many alerts it closed but how much of the analyst's attention it returned to the reading and the signing.
Which suggests a different way to think about what an AML function is buying when it invests here. It is not buying a faster analyst, and it is certainly not buying a replacement for one. It is buying a separation that the function has never been able to make on its own: the split between the labor of establishing facts and the responsibility of judging them. Those two things have been fused since the first monitoring system fired its first alert, and their fusion is why the cost of the negative finding has always been so punishing. Pull them apart, and the negative finding becomes cheap while the positive finding gets a sharper mind looking at it — the same analyst, on the same Tuesday, spending her day on the alerts that deserve a person rather than proving, several hundred times over, that most of them never did.
Discussion
No comments yet — start the conversation.