IA 360
AI Fundamentals

AI in Finance: Do Not Confuse Fraud, AML, and Risk

Fraud, anti-money-laundering alerts, and credit need different labels, errors, and controls. Reconstruct the decision.

Admin IA360 4 min read AI-generated Leer en español
AI in Finance: Do Not Confuse Fraud, AML, and Risk

On July 30, 2026, “AI for detecting financial risk” can refer to problems that barely resemble one another: blocking a transaction in milliseconds, prioritizing an anti-money-laundering alert, estimating default probability, or projecting portfolio losses. Grouping them under anomaly detection hides different data, consequences, and rules. The useful skill is to reconstruct each system as a decision: which event it predicts, when it must act, which error harms whom, and how it is validated after deployment.

Fraud, money laundering, and credit do not share a label

In payment fraud, the unit may be one transaction and the action may be to authorize, decline, or request verification. The answer arrives later and is not always clean: a chargeback, complaint, or investigation can change the label. In anti-money laundering, an alert does not declare guilt; it routes a case into analysis within a regulated process. In credit, the outcome may be default within a defined horizon, and the decision affects access, price, or limit.

This separation prevents a common error: assuming money laundering, terrorist financing, tax evasion, and card fraud are equivalent “anomalous patterns.” Behavior, frequency, available information, and responsible authority all change. An anomaly model can find something unusual without knowing that it is unlawful. A customer moving country or a seasonal business can also depart from its history.

Before selecting an algorithm, write a decision card: population, unit, horizon, label, action, cost of false positives and false negatives, and path for appeal. Then set a baseline using existing rules, human review, or the previous model. Only at that point does it make sense to ask whether a tree, network, or unsupervised detector improves the decision.

Fraud detection means learning with delay and an adversary

Positive cases are usually rare, so global accuracy misleads. Relevant measures include recall at a tolerable false-positive rate, alert precision, prevented economic loss, and friction imposed on legitimate transactions. Rates should be accompanied by absolute volumes: a decimal improvement may mean thousands of reviews. A threshold is not merely statistical; it distributes work, delay, and risk among customers, merchants, and analysts.

Time dominates evaluation. A random split may show a model the patterns of a campaign that had not yet appeared at the simulated production date. Testing should reproduce training on the past and deciding on a later future. Features must also be frozen to what was available at that moment. Using the outcome of a later investigation or chargeback would leak the answer.

The adversary adapts. Once a rule blocks one behavior, fraud moves to another; when verification is added, attempts and observed labels change. A Bank for International Settlements paper proposes a two-stage framework for [anomaly detection in high-value payment systems] and tests it with artificially manipulated transactions. It is a methodological demonstration for that scenario, not a universal performance rate for cards, transfers, or every attack.

Human review is not a free label either. Analysts see only cases routed into their queue, work under limited time, and may disagree. The system should record which evidence they consulted, when they overrode a recommendation, and which outcome emerged later. Sampling some low-scored transactions helps estimate what the filter misses; without that exploration, the model is judged on a set it selected for itself.

An anti-money-laundering alert begins an investigation

Rules capture known typologies and are easy to document, but they can generate many repeated alerts. Supervised learning needs reliable outcomes from earlier investigations; training only on cases the old system chose to review inherits its field of view. Unsupervised methods discover unusual behavior but can rarely label it a crime without context. A sensible architecture combines rules, learned signals, relationship networks, and expert judgment.

The network matters because splitting behavior across accounts or entities can make isolated transactions look normal. The BIS annual-report chapter on [AI and the monetary and financial system] discusses using payment-network patterns, customer information, and transactions against money laundering, as well as the challenge of combining data across jurisdictions. More data sharing is not automatically better: purpose, quality, access, and protection remain design requirements.

The final metric must follow the workflow. Beyond historical recall, track alerts per analyst, time to decision, duplicates, escalated cases, stability, and the ability to uncover new typologies. Fewer alerts may mean efficiency or blindness; more may mean coverage or noise. A report should show what happened to discarded cases and how investigation outcomes return to the system.

Credit risk requires calibration, fairness, and truthful reasons

A credit model estimates an outcome under observed conditions, not a person’s essence. Historical data contain earlier decisions: people who did not receive credit could not produce the same repayment history, while pricing, collection, and extension policies affect the outcome. This selection limits what can be learned. Validation needs time windows, segments, probability calibration, stress tests, and comparison with simple policies.

The European Banking Authority’s [guidelines on loan origination and monitoring] cover governance, creditworthiness assessment, pricing, collateral, and monitoring. Its [follow-up report on machine learning in IRB models] emphasizes prudent use within the internal-ratings framework. A predictive gain does not displace model governance or customer-protection rules.

In the United States, the Consumer Financial Protection Bureau’s circular on [adverse actions based on complex algorithms] says creditors must be able to provide specific and accurate principal reasons; technical opacity is not an excuse. An approximate explanation that does not reflect the factors actually used may be convenient for the model and false to the applicant.

In the European Union, the [Artificial Intelligence Act] lists systems used to evaluate natural persons’ creditworthiness or establish a credit score among high-risk uses, while excepting systems used for the purpose of detecting financial fraud. That classification shows why all “financial ML” should not be treated as one risk. The exact duty depends on intended use, actor, and the applicable legal timetable.

Managing the model matters as much as training it

A serious system has an owner, inventory, data and version documentation, limits of use, independent validation, approvals, monitoring, and a retirement plan. It records human overrides and fallback routes, checks that human reviewers have real information and authority, and rehearses missing data or supplier failure. A post hoc explanation does not replace understanding variables, transformations, and dependencies throughout the process.

When a third party supplies data, models, or infrastructure, responsibility does not disappear into the contract. An institution must know about version changes, coverage, latency, data retention, and exit conditions; reproduce critical tests with its own information; and keep an operational alternative. The best score loses value when nobody can explain why it changed or restore service after a dependency fails.

On April 17, 2026, the Federal Reserve, OCC, and FDIC issued [revised guidance on model risk management], replacing the previous United States guidance and emphasizing an approach tailored to an organization’s model-risk profile, size, and complexity. Sound development, validation, governance, and controls apply to the decision system, not merely the estimator’s code.

Monitoring should connect signal to consequence: input drift, prevalence changes, calibration, false positives, outcomes by group, operational workload, complaints, and incidents. When a measure crosses a limit, a planned action is needed: investigate, adjust the threshold, return to a safe version, or stop. To read any claim, ask which crime or risk is defined, what the label is, which action is automated, who bears each error, how an appeal works, and what evidence shows benefit beyond the sample. That is how to distinguish a statistical demonstration from a governable financial decision.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close