Skip to main content
Back to Blog
Compliance AI

Practical AML Anomaly Detection for Compliance: PIT to Production

September 202618 min read
AML anomaly detection for compliance

Anomaly detection uses machine learning to understand normal customer and transaction behavior, then identifies deviations that slip past rules-based systems. When done correctly, anomaly detection reduces overall volume of low-value alerts and enables investigators to prioritise alerts most likely to result in suspicious activity reports. It requires clean point-in-time data for best results, explainable outputs tied to known typologies, and ideally, a tested and documented pilot you can show regulators. Below, we cover the modelling decisions, data prep, and governance processes that help you build a working solution versus a costly exercise.

TL;DR:

  • Supervised models perform adequately when there is enough labelled suspicious activity, however unsupervised and hybrid models excel at identifying previously unseen laundering behaviour.
  • Develop point-in-time features carefully and ensure strict data discipline. Otherwise your back-test results will leak information and will not be trustworthy.
  • Combating extreme class imbalance involves resampling, synthesizing new data, and threshold adjustments to maintain reasonable false positive rates while still achieving high recall.

Explanation methods such as Shapley values are necessary to translate model scores into inspectable human-understandable reasons regulators and investigators can work with.

  • Successful AML anomaly detection requires case management system integration in order to ensure the feedback and alert flows allow for continuous model improvement, performance management, and compliance.

How To Build Smarter AML Detection Systems

POD Technology builds tailored artificial intelligence, machine learning, and scalable enterprise software for complicated banking systems and mission-critical solutions.

Learn More

Table of Contents

Core AI and ML approaches for AML anomaly detection

Rules-based monitoring identifies a suspicious transaction if it exceeds a predetermined threshold. Anomaly detection learns what normal behaviour looks like for a customer, peer group or network then identifies anomalies from that norm, including patterns no rule writer envisioned. Selecting from the many approaches available depends on the data you have, whether you have lots of labeled history, and what type of laundering scheme you're attempting to detect.

Supervised models can be effective when you have enough confirmed STRs to train against. Gradient boosted trees and logistic regression stay popular as they’re easy to retrain and explain to an auditor. They work well for understood typologies like structuring or swift movement of funds.

Unsupervised methods, like isolation forests, autoencoders and clustering don’t require labelled examples of fraud. They work best at surfacing genuinely new patterns, which is important because launderers innovate faster than any rules engine can keep up. The downside is a higher false positive rate until the model is refined based on investigator feedback.

Semi-supervised approaches incorporate elements of both: train the model on a small labelled dataset then utilize unsupervised signals to detect when data starts drifting away from the patterns seen in those labels. These methods generally represent the realistic compromise teams end up implementing when they have some labels available.

Graph neural networks (GNNs) take into account the account, counterparty and transaction network instead of analysing transactions independently. Research with privacy-preserving graph-ML pipelines has shown that graph-based ML learns network-level relationships that transaction-level models cannot detect, such as layering via multiple shell accounts. This applies primarily to organized laundering schemes rather than single-account outliers.

Many organisations take a middle-of-the-road approach: maintain existing rulesets as a safety net, augment them with ML scoring for prioritisation, and retire rules gradually once the ML model has been battle-tested.

  • Transaction monitoring: unusual patterns of transfer when compared to a customer’s own history are flagged by unsupervised or hybrid models.
  • Onboarding and KYC: supervised models rate new-customer risk with enriched identity and document data.
  • Sanctions screening: graph and entity-resolution methods identify attempts to disguise beneficial ownership through related entities.

Data and feature engineering essentials for reliable models

Trust your model as much as you trust your data pipeline. The number one reason AML models fail in production, even though they have strong back-test results, is due to data leakage from features incorporating information that isn't available at transaction time.

Point-in-time (PIT) feature construction addresses this issue by guaranteeing that each feature view contains only information available at scoring time, and not information that appeared later, like an account getting closed after an investigation completed. Correctly constructing PIT features with rigorous timestamp discipline in your feature store is what enables back-testing results to be trusted versus misleadingly optimistic.

  1. Aggregate canonical data sources — transaction logs, KYC and onboarding fields, sanctions and PEP list matches, device and IP signals, and historical alert or case activity.
  2. Normalize inconsistent formats across systems before calculating any features.
  3. Resolve entities — associate accounts, addresses and devices to the same underlying customer or network, as many typologies will only show themselves after entities have been linked.
  4. Cluster customers. Put customers with similar declared profiles such as industry, transaction volume, and geography into peer groups. This enables you to compare like with like when looking at behavioural baselines, rather than flagging all high-volume merchants as anomalous.
  5. Run reconciliation checks: ensure features are complete and fresh against source systems on a schedule. Alert when a data feed becomes stale or a field's null rate unexpectedly increases.

Guidance issued by Singapore’s MAS guidance on effective transaction monitoring controls highlights periodic back-testing of parameters and continuous data-quality checks as foundational requirements for effective controls. The same holds true for ML-based systems.

Tip: Log the date and time stamp of every feature's value when it's calculated. That's how you can discover leakage before some regulator tells you they found it.

Source DataTxns • KYC • PEPDevice • CasesPIT CutoffOnly data knownat scoring timeFeaturesPeer groupsVelocity • GraphModel ScoreAlert priorityExplainable outputLeakage boundaryFuture events must not cross this lineHistorical closure statusExcluded from training rowTrusted back-test and audit trail

Model selection and techniques to handle class imbalance

The number of confirmed money laundering cases is much smaller relative to the number of legitimate transactions. Often, the difference is many orders of magnitude. This scarcity drives nearly every modeling decision.

Algorithm choice is often a trade-off between explainability and capacity. Algorithms such as gradient-boosted trees and logistic regression are easier to explain to investigators and auditors. Algorithms like deep learning and GNNs can learn more complex patterns, but require additional effort if interpretable outputs are desired.

  • Resampling techniques, including SMOTE variants, rebalance training data without simply duplicating rare examples.
  • Synthetic graph augmentation creates realistic looking network activity around seed money laundering patterns to provide graph models additional training data.

Threshold tuning moves the decision boundary after training instead of during training. This allows the compliance team to decide how many alerts they can realistically handle.

  • Focal loss alleviates the problem of easy examples dominating the training process by down-weighing their contribution.

Temporal validation is just as important as class imbalance. Training and test data should always be split by time, never randomly. This ensures that you don't introduce lookahead bias, where your model appears predictive but has actually just memorised data collected after the time it’s meant to flag.

Interpretability methods, especially Shapley value attributions or feature-group attributions, convert a model’s numeric score into a story an investigator can act upon: what transactions, relationships or behavior changes triggered the alert. Without this step, an ML alert is no more useful than a black box, no matter how accurate.

Evaluation metrics and realistic baselines for AML models

Correct metrics determine whether your pilot is approved for production. False positive rate, precision and recall, alert-to-STR conversion rate, time-to-investigation, and serving latency against your SLA all require baseline measurements in order to establish improvement.

Legacy rules-based transaction monitoring scores very low on these metrics. According to industry research from SWIFT on AML and fraud detection, rule-based systems have false positive rates of between 70 and 95% on average, with rates of alert-to-STR conversion at only 2 to 10%. Measure any ML pilot against this baseline, not zero.

Typical rule-based transaction monitoring systems have false positive rates between 70–95% and alert-to-STR conversion rates of only 2–10%. If you have a pilot that significantly reduces false positives without impacting recall, you're already creating value above baseline.

Define acceptance criteria prior to pilot launch: an agreed amount to reduce false positives by, an acceptable floor for recall on confirmed historical cases, and a latency ceiling that your current case management SLA will tolerate. Run back-testing continuously, not only at launch, against live results with the same historical windows used during validation so you can detect drift sooner.

Explainability and compliance for AML anomaly detection

Regulators have been clear that experimentation with new AML tech is encouraged, so long as it’s properly reported. The FinCEN joint statement on innovative efforts to combat money laundering explicitly states that financial institutions will not necessarily be punished by examiners for piloting new approaches to BSA/AML compliance, as long as the pilot is prudently assessed and communicated with regulators.

It also makes clear that identifying weaknesses in current controls through piloting activity is not necessarily a criticism by supervisors. However, this is only the case if you record the findings and the pilot continues to comply with current requirements. The need to record is the cost of implementing ML in a regulated industry. You can’t avoid it.

Technical explainability and governance paperwork need to work together:

  • Tie model results to known typologies: all model results should link back to the laundering type they are intended to catch, rather than simply an anomaly score.
  • Keep model cards: note training data used, features selected, performance, and known limitations for each iteration of the model.
  • Maintain version history: record the date of each retrain, threshold adjustment, or feature update.

Record investigations: associate each alert with an investigator's justification and decision, creating an audit trail that demonstrates the model's real-world value.

As IBM states in its summary of AI in AML transaction monitoring, AI should be viewed as a tool to triage alerts and increase investigator efficiency versus a tool to automate human decision making. Explainability mapped to typologies is what allows human decision making to be audited.

Reduce-the-risk tip: engage your regulator early in the pilot design discussion, before you have results to share. Written engagement is a form of risk mitigation itself.

Deployment and MLOps for production AML systems

Transitioning from something validated to production forces you to answer questions your data science notebook never had to.

  1. Pick real-time or batch inference, depending on your scenario. Instant decisions are required for things like wire transfers with sub-second scoring, while batch can be used for periodic activities like behavioural reviews run overnight.
  2. Monitor for drift — both input drift, where feature distributions change as customers change behaviour, and model drift, where score distributions move away from their validated baseline.
  3. Watch alert volumes closely. Sudden increases or decreases often indicate a problem with incoming data well before they tell you anything about laundering.
  4. Consider serving format changes as model changes: transforming a model to ONNX or quantising it to increase speed can cause small shifts in outputs. Revalidate your model after serving format transformations instead of assuming they are the same as the original version.
  5. Close the feedback loop: feed investigator resolution outcomes back into the training pipeline on a regular cadence so your model can learn from both confirmed and dismissed alerts.

Having clear runbooks and named ownership in place across compliance, data engineering and the model team ensures this isn't an annual scramble.

Pilot-to-production checklist and common pitfalls

A skip-step pilot won't live through its first audit. Validate the scope is limited enough to validate cleanly before scaling anything, ensure PIT is used throughout your snapshot, and have a written validation plan with agreed upon thresholds for acceptance.

  • Define scope narrowly — one typology or customer segment per pilot, rather than the entire monitoring programme up-front.
  • Freeze a PIT data snapshot: without this step, your back-test numbers will be hopeful, not honest.
  • Test explainability: make sure you can explain every alert back to features and a named typology before presenting results to compliance executives.
  • Obtain regulatory approval of the pilot design, not just results, based upon regulator documentation expectations.
  • Set up post-deployment monitoring before go-live, not after the first drift incident.

Two alerts that should halt a rollout indefinitely are features leaking future knowledge and inexplicable increases in alert volume that no one can correlate to a business need. These are generally indicative of an issue with your data pipeline, not actual changes in customer behavior.

Responsibility should divide cleanly. Compliance would own typology definitions and sign-off. The data engineering team would own the pipeline and PIT discipline. The ML team would own model performance and explainability artefacts.

Expert Tip: always run your pilot’s explainability tests before you run its accuracy tests. If you can’t explain it, it won’t pass muster during the accuracy phase either.

PODTECH’s delivery approach for ML-driven anomaly detection

PODTECH’s AML anomaly detection engagements are divided into three phases: discovery, pilot and production. These services can be provided through a dedicated team or staff augmentation, depending on how your existing team is set up. During discovery, we map available data sources and define typology priorities. The pilot phase involves building PIT-reliable features, training and validating a model against agreed-upon acceptance criteria with compliance stakeholders, and documenting results for regulatory review. Finally, production solidifies the machine learning pipeline with model monitoring, drift detection and a defined model retraining schedule.

Controls persisted through each stage are point-in-time feature discipline, formalized feature engineering pipelines, and revalidation every time a model's serving format changes, including ONNX conversion or quantisation steps. Machine learning development applied by PODTECH follows these same engineering principles but within regulated, high-compliance environments. We're supported by 24/7 global support and decades of project delivery.

Integrating anomaly detection into existing AML infrastructure

An anomaly detection model outside your case management system will produce scores that no one intervenes on. Bidirectional integration is required: send alerts into the case management queue, and send investigator resolution back into the training pipeline.

In practice, this means that the model’s output should go into the same alert queue investigators already see, including its risk score, the typology it hit, and ideally the individual features which triggered the alert, rather than being provided as a second feed that requires another login or manual export. Case management systems which allow API-based ingestion of alerts simplify this greatly over those which require batch file exports, since near-real-time scoring is pointless if the alerts are delayed in a nightly import queue.

Feedback upstream is equally important. When an investigator closes an alert as a false positive, or escalates it to an STR, that resolution has to propagate back to the model’s training pipeline as structured, labeled data. Without that feedback loop, the model never learns from the decisions it’s supposed to be assisting with, and its false positive rate will inevitably regress toward the legacy baseline.

Legacy transaction monitoring systems and next-gen ML scoring engines frequently require an intermediary layer to normalize disparate data models and alert taxonomies. Purpose-built enterprise automation tooling for this type of workflow management is far more robust than custom point-to-point integrations, which create exponential maintenance burdens with each modification to either system.

Data privacy and security in AML anomaly detection

AML solutions manage some of the most sensitive information within an organization: complete transactional data, identity documents, device fingerprints and, more recently, behavioral signals. Such aggregation of information elevates risk associated with every security decision.

Training data as well as model results should only be accessible based on role-appropriate permissions. For example, an investigator should only see the alerts and associated case information that are placed in their queue. The investigator should not have direct access to the underlying feature store. Encryption in transit and in volume as well as at-rest are table stakes. Third party vendors that will have access to this data, whether that be a cloud provider or model serving platform, should undergo separate security evaluations prior to integration.

With institutions increasingly interested in sharing signals between entities without having direct access to underlying customer records, privacy-preserving collaborative approaches will become increasingly important. Research into privacy-preserving graph-based AML pipelines has been published that shows financial institutions can preserve network-level relationships, potentially indicating laundering rings across multiple banks, while minimizing how much raw data is shared between parties. This technology is still a research topic, not a regulatory standard, but any agreements between institutions to share data would still need to be legally vetted.

Retention policies are important as well: only storing training data and model logs for as long as you actually need them for validation and auditing will help minimize the attack surface while still providing the documentation you'll want ready during a pilot review.

Graph-based approaches appear to be transitioning to production more rapidly than most AML techniques. State-of-the-art research on Fourier-based contrastive learning for AML has demonstrated superior F1 scores on benchmark datasets when compared with several baselines, indicating representation-learning methods designed for other industries may be starting to gain real-world traction in the detection of laundering.

Privacy-preserving and federated approaches are another obvious direction, spurred on by the same pressures that fuel information-sharing consortia: laundering networks cut across institutions, but data privacy regulations restrict the flow of raw data between them. Methods that allow models to learn from patterns present in multiple institutions, without centralising that data, will likely become increasingly important as these networks mature.

Look for further convergence between anomaly detection and case management. The catalyst will be less the development of any particular new algorithm and more compliance officers' continuing obsession with auditability. An algorithm that results in a marginally higher score but a poorer explanation represents a compliance step backwards, and vendors selling to this market are responding by building explainability into their models from the ground up rather than trying to tack it on at the end.

Author perspective: pragmatic view on AI adoption in AML

When you hear about AI scoring a realistic victory in the AML world, it doesn’t automatically mean a completely hands-off detection system. Instead, think of AI as augmentation: smarter triage, less investigator hours wasted on false alerts, alerts that come with a reason attached. Teams who rush to build the fanciest model architecture without first addressing problems in their underlying data pipeline will typically build a black box that fails its first regulatory inspection. Nail down your point-in-time data discipline, bake explainability into your solution from day one, and document your pilot as if a regulator is going to take it home and read it because, chances are, one will.

— Harry

How PODTECH can help with AML anomaly detection

Creating a defensible AML anomaly detection system requires getting the data pipeline, model, and audit trail correct simultaneously. That’s a classic example of the kind of layered engineering PODTECH was designed to solve. Instead of a turnkey AI vendor, PODTECH functions as a delivery partner specifically for the pilot. We’ll leverage PIT, feature engineering, and revalidation best practices discussed throughout this article during a scoped engagement with your compliance team.

  • Develop machine learning: work on model creation, training and validation for anomaly detection pilots centered around your current typology priorities.
  • Enterprise automation: connect your model output to case management workflow such that alerts are routed directly to investigators without manual handoffs.
  • Systems integration: tie new ML scoring engines to existing legacy transaction monitoring systems and data sources.

Looking to pilot? Get an independent technical review of your existing data pipeline? Augment your team to develop a model in-house? PODTECH’s machine learning development team can help.

Visit Podtech to scope a pilot or discuss your current AML monitoring setup.

Disclaimer: This article provides general information. It is not intended as a substitute for nor should it be taken as financial advice from a qualified professional. Please consult your financial professional about your individual situation before making financial decisions.

Sources

FAQ

What are the five red flags in AML?

Examples of red flags are structuring activity, or breaking down a transaction to avoid hitting a reportable threshold; unexplained wealth or income that is out of character for a customer's risk profile; unusually quick transfers between accounts; transactions lacking an apparent business purpose; and transactions linked to high-risk locations or politically exposed persons. Anomaly detection looks for red flags such as these, even if no individual activity is over a certain amount.

Can you give me an example of anomaly detection?

A typical example would be a customer whose behavior suddenly changes from their own historical baseline. For example, a small retail account that starts to receive a few large international transfers over a few days. If we trained an unsupervised model on that customer's peer group, it would detect the change as anomalous without writing a rule specific to that situation.

What are three stages of AML?

The three stages are placement, where illicit funds enter the financial system; layering, where funds move through numerous transactions or accounts to distance them from their source; and integration, where the funds are absorbed into the economy and appear legitimate. Anomaly detection is especially helpful during layering, because graph-based models can observe the movement of funds between accounts, unlike rules-based systems that view each account separately.

What triggers AML checks?

AML checks are usually triggered by transactions meeting certain thresholds, behavior that's suspicious compared to a customer's profile, matches to sanctions or PEP lists, or risk indicators uncovered during onboarding, KYC, or continuous diligence. With ML, an additional method to trigger a check is having a model score surpass a predetermined threshold that maps to a documented typology that can be investigated.