There is a number that sits at the centre of most AML operations, rarely challenged and almost never examined for what it actually means. Depending on the institution, somewhere between 90 and 99 percent of transaction monitoring alerts turn out to be false positives. Industry estimates, vendor benchmarks and regulatory commentary have converged on this range for years, and the vendors that sell transaction monitoring infrastructure have documented the same burden across their client portfolios.
The number is not contested. What is strange is how calmly the industry has absorbed it. Think about what that figure actually says. A function staffed with trained investigators, running models that took months to build and tune, generating hundreds or thousands of alerts per week, and somewhere between 90 and 99 percent of what it produces is noise. In almost any other operational context, that ratio would trigger a root-and-branch redesign. In AML, it gets treated as the cost of doing business. The acceptance reveals something important about how the function has been allowed to define its own success.
FATF's mutual-evaluation methodology pairs technical compliance with an effectiveness assessment built on eleven Immediate Outcomes introduced precisely to test whether the regime works, not merely whether the rules exist. Yet that assessment still rests heavily on the outputs of the regime itself: reports filed, prosecutions brought, assets confiscated. The circularity has been noted by scholars for nearly two decades, most directly by Levi and Reuter (2006) and later Levi, Reuter and Halliday (2018), and it has never really been resolved. A system that measures its own activity as evidence of its effectiveness will always find ways to appear productive, regardless of whether it is. Transaction monitoring functions have inherited exactly that logic. The health of the function gets assessed through queue volume, alert throughput, and SAR filing rates. These numbers go into dashboards, board packs, and regulatory conversations. They become the evidence that the system is working.
But a high alert volume combined with a high closure rate tells you almost nothing about whether genuine financial crime is being detected. It indicates that the system is generating flags and that analysts are processing them. Those are different things, and treating them as equivalent is where the trouble begins.
A model calibrated to maximise coverage floods the queue with weak signals. Analysts facing a queue that they cannot fully work on apply less scrutiny to each case. Real risk gets buried in the noise. The SAR filing rate remains stable, sustained by defensive reporting, the well-documented pattern in which institutions file to protect themselves rather than because detection has improved.
That is not an edge case or a management failure at a poorly run institution. It is the ordinary operating condition of most transaction monitoring functions, sustained by a measurement framework that cannot distinguish between the two.
Here is a more concrete version of the problem. Take, for example, a mid-sized retail bank with 12 investigators reviewing 800 alerts per week. At 20 minutes per case, already an optimistic assumption once you account for SAR preparation, quality review, and system navigation, the queue consumes over 260 analyst hours before a single escalation is written. The economics alone make a meaningful review structurally impossible. What gets produced is throughput. Not judgment.
The cognitive consequences of sustained throughput under those conditions are well documented, just not yet in financial services. Research on clinical decision support systems showed that clinicians managing high-volume alert environments began overriding or ignoring warnings at rates that in high-volume settings reached well over 90 percent, with override rates across the literature commonly reported from roughly half to the high nineties (van der Sijs et al., 2006; Ancker et al., 2017). When the same stimulus produces no meaningful outcome hundreds of times in a row, the brain recalibrates its threshold for what counts as a signal. That is not a moral failing. It is how human attention actually works under sustained low-signal load.
The AML industry has not seriously reckoned with this. The design assumption embedded in most transaction monitoring operations is that analysts are stable review machines. Route a case to them, and they will assess it with the same quality of attention regardless of what the preceding two hundred cases looked like. That assumption is wrong. The systems built on it are built on a fiction about human cognition that the adjacent medical literature dismantled two decades ago.
The direct causal link between alert fatigue and specific missed suspicious activity reports is hard to establish because the counterfactual is unobservable. But the cognitive pathway is not speculative. It is the same mechanism that produces clinician desensitisation, and there is no principled reason it would operate differently when the alert is a transaction flag rather than a clinical warning.
A model cannot file a SAR. It cannot weigh narrative context, assess the plausibility of an explanation, or make a judgement call that withstands regulatory scrutiny. Those things require a human analyst with an intact capacity for careful attention. An oversaturated queue systematically erodes that capacity, which means it erodes the thing the entire function depends on. It does so invisibly, because throughput metrics never capture it.
This is the point where the argument is most likely to be misread, so it is worth being precise.
The fraud detection world has largely moved toward precision as the primary production objective. A model that generates too many false positives degrades operations faster than one with slightly lower recall, because operations are where actual decisions are made. That logic is sound, and much of it transfers to AML. But not entirely. High recall remains non-negotiable in certain contexts: elevated-risk customer segments, specific cross-border corridors, jurisdictions under heightened scrutiny where the cost of a missed pattern is regulatory and reputational rather than just financial. A compliance officer reading this will rightly observe that a regulator examining your function after a failure will not accept "we were optimising for analyst capacity" as a sufficient explanation for a missed high-risk pattern.
The argument is not for indiscriminately reducing alerts. It is for calibrating alert generation to what the operation can realistically support, so that cases reaching review actually receive the quality of attention that enables detection. That is a meaningfully different position, and conflating the two hands critics an easy rebuttal that the underlying argument does not deserve.
I have seen a similar miscalibration in collections analytics. A payment arrangement model tuned too aggressively buries the operations team in weak cases, slows resolution across the portfolio, and ultimately produces worse outcomes even for the accounts that most needed intervention. The fix was not more sophisticated modelling. It provided a clearer definition of what the model was supposed to produce for downstream operations: a workable decision, not a probability score. AML monitoring has the same structural problem, at higher stakes, and with considerably less institutional appetite for naming it.
A more useful frame for the function's effectiveness starts with the quality of decisions, not the volume of activity.
Of the cases that proceeded to SAR filing, how many originated from model alerts as opposed to internal referrals, third-party intelligence, or analyst-initiated review? If the model is not the primary source of confirmed detections, that is important information. It tells you whether the alert queue is contributing to the function or simply consuming its resource while the actual detection happens elsewhere.
What is the false positive rate disaggregated by alert type, customer segment, and transaction channel? A single aggregate figure conceals operationally significant variation. A rule that performs adequately on domestic retail transactions may be generating near-total noise in cross-border SME flows. Treating them as a single number makes both appear better or worse than they actually are, and it makes targeted intervention impossible.
What happens to analyst output quality as queue volume increases over sustained periods? This is almost never tracked because the case management systems that could capture it are not configured to do so. But the institutions that do monitor escalation rates, narrative depth, and case quality across rolling periods tend to find degradation that tracks queue volume almost directly. Not because the analysts have become worse at their jobs, but because the conditions have changed in ways the throughput metrics never reflect.
And underneath all three sits a simpler question that rarely gets asked in board packs: does anyone in the function actually know, on a case-by-case basis, which rules are earning their place in the queue? Most institutions could answer the first three questions by pulling data. This one requires someone willing to say, out loud, that a rule the model has run on for years might be worth nothing.
For SAR quality itself, which is difficult to measure directly, useful proxies include the rate at which filed SARs generate law-enforcement information requests, the depth of escalation within individual cases, regulatory feedback on submission quality, and narrative completeness scoring. None of these is perfect. All of them sit closer to the actual objective than the closure rate within SLA.
When alert volumes become unsustainable, the default institutional response is to tune the model: lower a threshold, add a suppression rule, retrain on a new data window. Sometimes that helps at the margins. Mostly, it relocates the pressure without resolving it.
The model is being asked to carry out a function that requires a much broader architecture to operate properly. Entity resolution, customer risk tiering, behavioural baseline modelling, analyst feedback loops, alert routing logic, case management workflow, and monitoring of the monitoring system itself: all of these have to hold together. When any one of them is weak, the model's output degrades in production regardless of how well it performed in development.
Entity resolution carries more weight in that list than it is usually given, because a meaningful share of false positives are network-context failures rather than model failures: an alert fires because the system is scoring an isolated transaction, when the question that matters is who the counterparties resolve to and what the surrounding network looks like. Graph-level context resolved entities, shared attributes, and transaction paths are precisely what separates a coincidental cluster from a real mule ring, and no threshold tuned on isolated transactions can recover that distinction.
The feedback loop deserves particular attention because it is the most consistently broken component. Most AML models operate in environments where structured outcome data never flows back to the model in usable form. Either because the case management system does not capture outcomes at the granularity the model needs, or because the timeline between alert and confirmed outcome stretches well beyond any practical retraining window. The result is a model that cannot learn from its own production errors at a useful speed. It improves in backtests but stays static in deployment, which is exactly the opposite of what the operating environment requires.
That gap between offline performance and production behaviour is one of the more reliable findings across applied ML domains. AML is not exempt from it, and no amount of model sophistication can resolve it without investment in the feedback architecture surrounding the model.
Regulators assess AML functions on the quality of controls, not just their existence. A function generating 900 alerts per week with 95 percent closed within SLA looks strong in a management information report. It may look quite different to a skilled examiner who asks to see the case notes, review the rationale for closures, and trace the evidentiary chain from alert to suspicious activity.
An analyst who has worked through 400 weak alerts this week, with 20 minutes available for each remaining case in a queue of 40, is not positioned to deliver the quality of documented reasoning a regulator expects. The control exists on paper. The control quality is something else entirely.
What is odd is how rarely this is framed as a compliance risk in its own right, given how directly it follows from the operating conditions under which most AML functions actually run. The compliance argument for fixing alert fatigue is at least as strong as the operational one. It just requires naming something that the function has a structural incentive not to name.
Before the next model tuning cycle, the more productive conversation starts with the function's operating assumptions and can largely be reduced to two questions. What percentage of alert-generating rules have never produced a SAR, across their entire production lifetime? Some zero-hit rules earn their place as required coverage a typology the firm is expected to monitor whatever the hit rate, where switching it off is itself the regulatory finding. But many do not, and most functions never separate the two. A rule that has run for years without contributing to a detection and answers to no coverage requirement is not a safety net; it is a noise generator consuming capacity that could be directed to cases that actually matter. Most functions do not know which of their rules fall on which side of that line. They should.
What queue size allows real analytical scrutiny, given actual investigator headcount and realistic case complexity, and how quickly does the outcome of that scrutiny make its way back into the model? Working backwards from analyst capacity to the required precision is a more honest calibration exercise than working forwards from a model threshold to whatever queue size results, and then managing the downstream consequences. The retraining lag is part of the same question: a detection system that cannot learn from its own production behaviour at any meaningful speed will always be slightly behind the patterns it is supposed to catch, and that gap quietly compounds over time, regardless of how well the queue itself is sized.
The goal is not to generate alerts, nor to close them quickly, nor to file a defensible volume of SARs each quarter and produce a report demonstrating that the function was active. The goal is to identify genuine financial crime while intervention is still possible and the information is still actionable. Everything else, the models, the rules, the thresholds, the queues, the dashboards, is infrastructure in service of that objective. When the infrastructure starts consuming the objective, which is precisely what happens when alert fatigue takes hold inside an operation measuring the wrong things, the function has lost the thread of its own purpose.
What makes this uncomfortable to address is that the measurement framework that produces the problem is also the one most institutions present to regulators as evidence of control quality. Changing it means acknowledging, at least internally, that what has been presented as a working system may be a busy one. Those are different things. The gap between them is where most of the risk actually lives.