An AI Mission for Banking: Dispute Resolution

Banks grade their dispute operations on whether they reached the right answer. The losses almost never come from the answer — they come from the case that was right and late.
A dispute arrives on a Friday afternoon as a short, unhappy sentence in a chat window: a cardholder does not recognize a charge. Within minutes it becomes a case record, and the moment it does, several clocks start running at once, each belonging to a different authority and each indifferent to the others. There is the window in which the bank owes the customer an acknowledgment, and a separate one governing provisional credit. There is the card network's own calendar for filing a chargeback, and behind it the merchant's window to represent, and behind that a further set of timers for pre-arbitration and arbitration if the thing goes the distance. There are internal service commitments that are looser than all of these but that the contact center will be measured against anyway. None of these clocks is visible in the chat window. None of them appears on the analyst's screen as a countdown. They exist as institutional knowledge distributed across a procedures document, a network rulebook, a compliance memo, and the particular experience of whoever happens to be handling the case.
Now watch what actually goes wrong. The analyst who picks up the case reads it correctly, gathers the transaction detail, and forms an entirely defensible view of the merits. Then the case waits — for a document the customer has been asked to provide, for a merchant response that has not arrived, for a fraud team's second look, for the analyst themselves, who has forty other cases and is triaging by whatever surfaced most recently. Somewhere in that waiting, one of the clocks expires. The bank does not lose because it decided wrongly. It loses because a correct decision arrived after the only window in which the decision could be acted on, and by then the outcome had already been determined by the calendar rather than by the facts.
The adjudication was rarely the hard part
If you sit with a dispute operation for a week, the striking thing is how seldom anyone is genuinely confused about the merits. The overwhelming majority of disputes fall into shapes the team has seen thousands of times: a duplicate charge, a subscription the customer thought they had canceled, a hotel incidental that posted late, a card-not-present transaction from a device the customer has never used. Experienced analysts resolve these almost reflexively, and the ones that are genuinely ambiguous — the ones where reasonable people would disagree about liability — are a small and identifiable minority. This is why the reflex to point AI at adjudication has a strange quality to it. It is aiming intelligence at the part of the process where the institution already has plenty, and where being faster changes very little, because the case was never stuck on the decision.
What the case is stuck on is state. At any given moment a dispute is in some position relative to each of the clocks that governs it, and that position is not recorded anywhere as a fact. It has to be reconstructed: which stage the case has reached, which evidence has been requested and which has actually come back, whether the network filing has been made, whether the representment landed and started a new timer, whether the provisional-credit posture is still appropriate given what has been learned since. A dispute case management system stores dates, and storing a date is not the same as understanding what the date implies. The system will faithfully tell you when the case was opened; it will not tell you that this particular case, in this particular posture, with this particular missing document, is three business days from a network deadline that nobody has looked at.
So the operational picture in most banks is a queue, and a queue is a fundamentally poor instrument for managing overlapping deadlines. Queues are sorted by something — age, priority flag, assignment — and whatever they are sorted by is a proxy for urgency rather than urgency itself. The oldest case is not necessarily the most exposed; a young case that has just triggered a short network window can be far closer to a hard edge than one that has been sitting benignly for weeks waiting on a customer who is never going to respond. Analysts develop instincts for this and the good ones are remarkable at it, but instinct does not scale across volume spikes, does not survive turnover, and does not work at all on the cases that never surface because nothing prompted anyone to look at them. The failure mode is not carelessness. It is that the information required to prioritize correctly is spread across systems and rulebooks that no human can hold simultaneously for hundreds of open cases.
Alerts notice; they do not track
The obvious objection is that this is what alerting is for, and every dispute platform has some version of it — a report of cases approaching a threshold, a colored flag, a daily extract that lands in someone's inbox. These help, and they leave the underlying problem almost exactly where they found it, for the same reason a dashboard never fixed a plant floor. An alert is a message that a condition was met; it still requires a person to receive it, interpret it against the specifics of the case, work out what the next action actually is, and then perform that action across whichever systems are involved. The alert has moved the work forward by one notification and left every subsequent step on the human. When the queue is under strain, the alerts pile up alongside the cases, and the team develops the entirely rational habit of ignoring a channel that cries wolf all day.
It is worth being honest that much of what is currently marketed as agentic AI for operations of this kind does not close that gap either. Gartner has predicted that over forty percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, inadequate risk controls, and what the firm calls "agent washing" — existing rule engines and workflow tools relabeled without any change in what they can actually do unattended. A rules engine that fires when a date approaches is a louder alert, not a different capability. It can only act on the deadline conditions someone thought to encode, in the case states someone thought to anticipate, and disputes generate exception states constantly: the partially received evidence, the representment that arrives with new information, the customer who calls back mid-cycle and changes the claim, the case that turns out to be one of six from the same merchant.
Put the mission on the clocks, not on the verdict
The higher-value design points intelligence somewhere less glamorous than adjudication: at continuous awareness of where every open case stands against every clock that governs it. That means something running persistently rather than on a schedule, holding the rules as knowledge rather than as hard-coded branches, reading each case's actual state from the systems that hold it, and reasoning about exposure — which cases are approaching an edge, which are blocked and on what, which have quietly changed posture because a document arrived overnight or a merchant responded. From there the work becomes tractable: chase the outstanding evidence, assemble the representment package from what the systems already contain, move a case forward when its path is unambiguous, and put in front of a human the small set of cases where the judgment is genuinely contested or where the money and the customer relationship warrant a person's name on the decision.
This is the shape that the autonomous enterprise body of work describes across operational functions generally, and disputes are an unusually clean instance of it, because the governing constraints are explicit, externally imposed, and unforgiving. It is also why StudioX frames capabilities of this kind as AI Missions rather than as assistants: a mission is a standing responsibility with an owner, not a prompt that answers when asked. Specialist agents watch the intake, the evidence, the network filings, and the customer communications; a reasoning core holds the case in its full context rather than as a row in a table; observations accumulate so the state of a case is something the system knows continuously rather than something an analyst reconstructs each morning; and human-in-the-loop sits where it belongs, on contested liability and on exceptions, not on the mechanical business of noticing that a timer is running out.
The reframing worth carrying is that a dispute is not really a decision with a deadline attached. It is a position against a set of clocks, and the decision is one event inside it. Banks that see it the first way will keep investing in accuracy they already have, and will keep absorbing losses that no accuracy metric explains, because the number that governs their outcomes is not how often they were right but how much of their portfolio was sitting somewhere unwatched when a window closed. The measure that matters is aggregate clock exposure — how many open cases are approaching an edge that nobody is currently looking at — and the institutions that start managing that number will find that most of what they used to call a dispute backlog was never a backlog of decisions at all.
Discussion
No comments yet — start the conversation.