
Before you put an AI model in the operations queue
Say your disputes team opens 300 cases a day, and an analyst spends the first few minutes of each one reading the cardholder's message and the transaction record to work out which reason code applies and which queue the case belongs in. A vendor shows you a model doing that step in two seconds. The demo works, because demos are picked for that. What a demo cannot tell you is how the model behaves in its second month.
Classifying, routing and summarizing cases is repetitive reading work, and current language models read quickly and reasonably well. We think they belong in some operations queues. An operations queue also carries money and scheme deadlines, and a model that misroutes a case does it at the same speed it routes the right ones. Most of the work of using one well happens before anybody signs a contract.
Choose one decision and write it down
Start narrower than feels necessary. "Help the disputes team" is a project. "Assign each new dispute to one of eight categories, using the cardholder's message and the transaction record" is a decision you can test. Write down the inputs the model sees, the answers it is allowed to give, and what a wrong answer costs in each direction. Sending a fraud claim to the general service queue can mean a missed scheme deadline. Sending a simple "item not received" case to the fraud team wastes an hour of a senior analyst's day. Those two errors are not worth the same, and your test should count them separately.
There is a useful line between a model that prepares work and one that decides it. Routing a case, suggesting a reason code or drafting a summary for the analyst is preparation, because a person still makes the decision that moves money or closes the case. Approving a refund or closing a fraud alert is the decision itself. Start with preparation, which is easier to test and easier to undo.
Check whether you need a model at all. A fair share of routing can usually be done with rules in the case management system you already own: a keyword in the message, or the transaction type. We made the argument for trying things in that order in three kinds of automation. If a rule sends most cases to the right place, the model only has to deal with the remainder, which is a smaller and cheaper problem.
Measure it against your own people first
Before the model acts on anything, run it in shadow. It sees each live case and records an answer while the analysts work as normal. Nobody reads the model's output until the comparison. A few weeks of shadow running on the real queue tells you more than any accuracy figure from a vendor's test set, because it includes your customers' spelling and whatever arrived on the Monday after a public holiday.
Then measure your analysts against each other. Take a couple of hundred closed cases and have two experienced people label them independently. They will probably disagree more than you expect, and that disagreement rate is the realistic ceiling for any comparison you make afterwards. If two of your best analysts agree on 85 percent of cases, a model that agrees with them 80 percent of the time is close to human performance. If your people agree on 95 percent, the same model is well short. Both figures are illustrative; the point is to know yours before you read the vendor's.
Break the errors down by category. An overall agreement rate hides the pattern that matters. Models tend to do well on the common categories and worse on the rare ones, and in disputes the rare ones are often the expensive ones.
Set the threshold, then staff below it
Most tools return a confidence score with each answer. Pick a threshold where the model's accuracy above it matches your human baseline, and send everything below it to a person. That share is the number to plan around. If 30 percent of cases fall below the threshold, you still need analysts for 30 percent of the volume. Those cases will be the harder ones, so they take longer than today's average.
Keep reviewing above the threshold too. A small random sample of confident answers should go to an analyst every day, and someone should read the results weekly. That sample is how you find out accuracy has drifted. Drift has ordinary causes: a new product, a scheme rule change that alters reason codes, a merchant with an unusual descriptor, or the vendor updating the model underneath you without much notice. Keep a fixed set of labeled cases and re-run it whenever the prompt, the model version or your category list changes.
The share of cases that falls below the confidence threshold is your staffing plan, whatever the headline accuracy figure says.
What your sponsor bank will ask
Your sponsor bank's risk team will want to understand how the model is used, and so, eventually, will your own compliance function. The questions are predictable enough that you can answer them in writing before go-live.
-
Who approved it, and who can change it?
Name the owner of the model's use in operations, and put prompt and configuration changes under the same change control as any other production change. An analyst editing the prompt on a Friday afternoon is a production change.
-
What does it see, and where does that data go?
List the fields sent to the model and where the provider processes and stores them. Keep full card numbers and anything else you would not put in an unencrypted email out of the prompt.
-
Can you reconstruct a single decision?
For every case, record what the model received, what it returned, its confidence score, the model and prompt version, and whether a person overrode it. When a customer complains about how their dispute was handled, you need to show why it went where it went.
-
What happens when it is off?
The queue has to keep moving through a provider outage or a decision to switch the model off for a week. That means analysts who still know how to triage by hand, and a routing rule that sends everything to them.
Summaries need one extra check. Once analysts trust a model's case summary, they stop opening the underlying messages. A summary that drops the one line where the cardholder mentions recognizing the merchant will be believed anyway. Keep a link from every summary to the source documents, and put summaries into the daily review sample alongside the classifications. If nobody on the team has read a raw case file in a month, your sample is too small.
Discuss your operations
