The exam was going well – until the second question.
The first question was easy: Walk us through how this customer’s risk rating was determined. The BSA officer clicked into the shiny new AI-driven risk engine, pointed at the score – 87, high risk – and explained the enhanced due diligence it had triggered. A textbook answer.
The second question was seven words long. “Why is this customer an 87?”
Silence. The officer started clicking around. The vendor’s documentation described the model’s architecture in general terms. The scoring factors were listed but not weighted; the weights were listed but not for this customer; Another customer in the portfolio – same industry, same geography, similar volumes – was a 52, and nobody in the room, including eventually the vendor on a hastily arranged call, could say why. The examiner made a note. That note became a matter requiring attention. It became 18 months of remediation for a system that had, in every statistical sense, been performing fine.
Why FCC Remains AI’s Toughest Challenge
Every industry talks about responsible AI. Financial crime compliance (FCC) has three features that make it unusually unforgiving.
First, the outputs are accusations. A monitoring alert, a risk score, a SAR – each is a step toward asserting that a specific human being may be committing a crime. That assertion enters government databases, can surface in law enforcement investigations, and follows people. A misfiring e-commerce recommendation engine shows someone the wrong pair of shoes. A misfiring compliance model puts an innocent person’s name in a file.
Second, the decisions must be reconstructible years later. A SAR filed today may be questioned by an examiner in 2028 or subpoenaed in 2030. “We have since upgraded the model” is not an answer. The institution must be able to back the decision with the relevant logic as it was made.
Third, the base rates are brutal. Actual laundering is rare within any customer population, meaning even a high-performing model produces more false positives. As a result, errors are the ordinary case, not the exception, and the pattern of who gets incorrectly flagged can itself become a compliance risk. If your alerts concentrate on customers with foreign names, remittance corridors, or cash-intensive small businesses at rates the underlying risk cannot justify, you have not built a monitoring system but a fair-lending problem with extra steps.
The Three Failure Modes
It helps to be clear about what actually goes wrong because the mitigations differ.
Opacity: This is the risk-engine story described above – a decision that cannot be explained at the level of the individual case. The fix is architectural, not cosmetic. It means choosing interpretable models when the stakes are high, requiring decision-specific factor attribution rather than marketing-level claims of ‘explainability’, and capturing the rationale at the moment the decision is made.
Any officer can run a useful test today: pick a random alert, ask your team to explain the score to you as if you were the customer’s attorney. Measure how long it takes. If the answer is “we’d need to ask the vendor,” you already have the finding; it just has not been documented yet.
Bias: It is quieter, no one intends it, and it arrives through proxies. Occupation codes, geography, transaction corridors, and name-matching algorithms that perform worse on non-Anglophone names all smuggle protected characteristics into models that never see a protected field. The mitigation is measurement: regular, documented disparity testing across customer segments with thresholds that trigger investigation. This must be paired with the harder organizational step of deciding in advance who reviews the results and what they are empowered to do. A bias metric nobody owns is a liability memo the institution wrote to itself.
Hallucination: The newest failure mode, and in evidence chains, the most dangerous. A language model asked to summarize a case will sometimes produce a confident sentence that is simply false – a transaction that never occurred, an adverse-media hit that belongs to a different person with a similar name, a corporate relationship inferred from nothing. In a marketing draft, this is embarrassing. In a SAR narrative, it is a false statement in a federal filing. In an evidence chain, it can metastasize: the hallucinated ‘fact’ gets cited by the next case that references this one, and by the third citation, it has the patina of an established finding. One institution’s QA team traced a phantom shell-company connection through four case files before finding its origin – a single generated sentence, never verified, 11 months earlier.
An unverified sentence in an evidence chain does not stay one sentence. It becomes a precedent.
The mitigation here is mechanical and non-negotiable: no generated factual claim enters a case record without a citation to a source system, and the citation must be clickable, checkable, and checked. Zero unverified facts in filings – not low hallucination rates – enforced by grounding requirements in the tooling and sampled verification in QA. Institutions that treat this as a model-quality metric miss the point. It is an evidence-integrity control – the same category as chain of custody.
What to Build Instead of Excuses
The pattern across all three failure modes is the same, and it is the actual content of the phrase ‘responsible AI’ once the conference-keynote varnish is off. You are not being asked to prove your models are perfect. You are being asked to prove you know how they fail, that you look for those failures, and that someone answerable acts on what you find.
That translates into a short, concrete build list – an inventory of every AI system touching a compliance decision. Most institutions discover their inventory is longer than they thought, once the vendor tools with ‘embedded AI features’ are counted honestly. A per-decision explanation requirement, graded by stakes. Disparity testing with named owners. Grounding and citation gates on anything generative that touches an evidence chain. And a written account, in policy, of which decisions are reserved for humans – the accusatory ones, the filed ones. Two external frameworks can do scaffold duty here: NIST’s AI RMF for the structure, and the EU AI Act’s high-risk obligations – taking effect from August 2026 – as a preview of the statutory floor for any institution with European reach.
None of this slows a good program down; that is the part the skeptics have backwards. The institutions with real explainability tune faster because they can see what the model is doing. The ones that test for bias catch broken data feeds early, because disparity spikes are often data defects wearing a disguise. The ones with citation gates trust their copilots enough to expand them.
And the officer from the opening scene? The officer’s remediated risk engine now produces, for every score, a plain-English factor breakdown that sits in the customer file. Last month, an examiner asked the second question again, about a different customer. The answer took only 40 seconds. The examiner moved to the third question, which is the entire ambition of the discipline: making the second question boring.
How SG Analytics Helps
SG Analytics builds compliance AI that can answer ‘the second question’ without hesitation. We design interpretable-by-design risk models, generate per-decision explanations written into the case record, conduct disparity testing with owners and thresholds, and implement grounding gates that keep generated text out of evidence chains until every fact carries a checkable citation. Our investigators and data scientists work as one team, under ISO 27001 and SOC 2 Type II controls, ensuring governance is built in rather than bolted on.For institutions inheriting opaque vendor tools, we run independent AI audits – covering inventory, explainability assessments, and bias measurement – to build a clear remediation roadmap. Because we are not tied to any platform, we can explicitly state when a tool should be replaced if it cannot be made defensible.
