S
SHUBHAM CHATURVEDI
Guest
A few months ago, I was investigating an unexpected spike in AML false positives. Every dashboard looked healthy. Every ingestion job had completed successfully. Infrastructure monitoring showed no issues. Initially, I assumed something was wrong with the pipeline. After digging deeper, I realized the real issue was the data itself. An upstream application had unintentionally changed how one field was populated. Nothing crashed or failed validation. But one missing attribute was enough to throw customer profiles out of sync and undermine the detection logic we trusted.
After tracing the problem through the pipeline, I discovered that an upstream application had stopped reporting closed-account status correctly. My AML models continued treating those accounts as active, which distorted customer profiles, inflated transaction histories, and generated investigations for accounts that should no longer have been in scope.
I only found the root cause after working backward from the spike in false positives, validating field mappings one by one until I identified the upstream source. The pipeline itself wasn't broken. The data flowing through it was. That experience reinforced an important lesson: even the best detection models are only as reliable as the data they receive. Therefore it is safe to say that this is not a theoretical problem. It is a systemic vulnerability across financial crime compliance that a vast number of institutions encounter as systems and data evolve.
The industry is seeing this problem more often than many people realize. According to ComplyAdvantage's State of Financial Crime survey, 45% of compliance professionals name poor-quality, hidden data as the top barrier to effective risk detection. Separate insights from McKinsey & Company point to fragmented, legacy data as the major contributor behind the massive volume of false positives eating up investigators’ time.
I see firsthand how teams often spend weeks tweaking detection logic, only to discover later that the data itself was broken from the start. If we don't audit our incoming data as strictly as we write our logic the whole system becomes susceptible to failure.
In standard business intelligence pipelines, bad data might simply result in a skewed chart or minor metric discrepancy. However in an AML pipeline, it may actively prevent the software from catching compliance issues which may lead to regulatory penalties and legal issues.
Some of the most dangerous financial crime gaps stem from four persistent data failures:
One of the most difficult issues to detect is data drift. In my experience, engineers building an app or website are constantly pushing updates to support new business requirements. They may rename a field, change its format, or start sending null values instead of valid data.
What follows is this problem called Schema Mutation. The app team made their changes, started receiving weird data but just passed it along. The pipeline didn't break or fail, it just started writing garbage data into that column.
Data Drift is another common consequence: The structure is the same, but the actual data inside shifted. For example, a country field used to store full names (United States) but now stores codes (US). The logic (filter or lookup) on this field may miss this since the expected data was different.
Since the pipeline can still process the records without errors, these changes often go unnoticed.
Old, disconnected systems make it hard for banks to share data. Records for transactions, Know Your Customer (KYC) details, and risk profiles are housed in separate systems, often with inconsistent structures.
It can be hard to describe how bad the structure can get. In one system, all dates may be stored as alphanumeric strings, and in another, floating point numbers are used for Customer IDs, with vital AML fields missing. There may also be a case of a completely different risk score with a duplicate profile. Given how fragmented the data is, resolving profiles (the ability to tell whether or not "John Doe" and "J. Doe" are different people) is not even possible. In the end, this results in split profiles that conceal how a customer transacts, and causes other systems to miss suspicious behavior or generate an overwhelming amount of false alerts.
A technically perfect system can still fail if the semantic meaning of the data changes. Structuring detection logic requires rigorous time-window joins—such as checking whether a customer conducted multiple cash deposits across different branches within a rolling 24-hour window to evade Currency Transaction Report (CTR) thresholds. If the timezone offset is dropped, truncated, or improperly cast during ingestion, the window shifts. The data remains clean and structurally sound, but the order of events is lost, making the review worthless.
The most dangerous systemic risk is that AML pipelines heavily depend on upstream data sources owned by teams who have no idea their technical changes carry massive regulatory consequences. Because there is traditionally no shared ownership or mandatory validation between core banking engineers and compliance engineers, data integrity is broken before it even reaches the analytics layer. Pipelines quietly depend on upstream data sources that are completely unstable, undocumented, or modified without alert-impact simulations.
Over the years, AML programs have focused heavily on creating detection rules and meeting regulatory requirements. While those are important, I've found that strong compliance starts with something much simpler—trustworthy data.
As data engineers, we spend a lot of time making sure information moves correctly from source systems to downstream applications. If that data is incomplete, inaccurate, or changes unexpectedly, even the best detection rules can produce unreliable results. That's why data governance and data quality should be part of every AML program, not an afterthought.
The industry is also moving toward a more evidence-based approach to compliance. Instead of relying only on predefined rules or thresholds, organizations are expected to demonstrate that their monitoring systems are built on accurate, complete, and well-governed data. A rule is only as effective as the data it evaluates.
Regulators are no longer interested in seeing only the alerts your system generated. They also want to understand why those alerts were generated—or why they were missed. That means you should be able to trace every important data element from the source system, through each transformation, to the final AML model or alert. Having that level of data lineage not only helps during audits but also makes it much easier to investigate issues when something doesn't look right.
Fixing these vulnerabilities doesn't require a flashier detection rule or a more expensive black-box AI tool. It requires implementing precise, rigorous data governance and data engineering practices that treat data integrity as a hard regulatory requirement.
To safeguard an AML program against silent data corruption, institutions must build their technical strategy around these foundational pillars:
To solve the upstream disconnect, organizations must implement Data Contracts. A data contract is a formal, version-controlled agreement between upstream software engineers and downstream data consumers that defines the expected schema, SLA, and semantic meaning of the data.
Coupled with this, automated Data Quality Firewalls must be pushed directly to the ingestion layer. Instead of cleaning data downstream in a warehouse or lakehouse, data must be validated the second it lands:
Schema Validation: Strict checking to ensure required AML fields (such as Originator/Beneficiary routing numbers or legal names) are populated and conform to explicit structures like SWIFT or ISO 20022 messaging standards.
Value Constraint Checking: Automated checks for null values, anomalous zero balances, or invalid data types (e.g., catching when a critical identifier accidentally converts to a float).
Quarantine Routing: Any data that violates the contract is automatically goes to a separate folder immediately so it cannot break the main system.
For an AML framework to be truly defensible, data cannot just be stored; it must be auditable across time. Using modern, open table formats that support ACID transactions and Time-Travel queries is essential.
Snapshot Isolation & Time-Travel: This capability allows engineers to query the data exactly as it existed at a specific millisecond in the past. If an investigator or an auditor asks why an alert is fired on a specific date, you can rewind data to see the exact moment of that customer's profile and transactions at that exact moment.
Deterministic Lineage: Every transformation—from raw string data to enriched customer profiles—must be mapped. Metadata platforms must capture the exact version of the detection code and the exact snapshot ID of the data used, ensuring absolute reproducibility.
Before deploying or altering any critical compliance model, engineering and compliance teams must operate in lockstep through a disciplined validation process:
Deconstruct Intent: Technical teams must profile raw data sources and explicitly map fields to the regulatory basis of the rule (such as the Bank Secrecy Act, FinCEN guidance, or FINRA sales-practice rules) alongside the compliance analysts who actually work the alerts.
Edge-Case Unit Testing: Build modular transformations that are unit-tested specifically against dirty data: threshold boundaries, nulls, duplicates, and same-day reversals.
Historical Parallel Validation: When migrating systems or tuning thresholds, run old and new pipelines in parallel using a minimum of 90 days of historical data to compare outputs exactly.
Volume Guardrails: Establish baseline alert volumes. In my own deployments, I've used a drift threshold of roughly 20% against historical baselines as the trigger point for automated guardrails to halt a rollout—the exact number should be calibrated to each institution's own alert volatility, but the principle holds: production rollouts need an automatic circuit breaker before compliance operations are disrupted.
One lesson I have learned is that an AML program is only as reliable as the data behind it. You can spend weeks refining detection rules and adjusting thresholds, but if the underlying data is incomplete or inaccurate, the results will never be reliable.
For me, improving data governance isn't just about building better pipelines; it's about building confidence in the decisions those pipelines support. Validating data, tracking lineage, and monitoring for unexpected changes help ensure that AML models continue to perform as intended.
Before changing another rule, ask yourself a simple question: Can I trust the data feeding it? If the answer isn't a clear yes, that's where the investigation should begin.
As AML systems become more data-driven, I believe data quality will play an even bigger role in successful compliance programs. Strong detection starts with trustworthy data, and that's something every data engineering team can help deliver.
After tracing the problem through the pipeline, I discovered that an upstream application had stopped reporting closed-account status correctly. My AML models continued treating those accounts as active, which distorted customer profiles, inflated transaction histories, and generated investigations for accounts that should no longer have been in scope.
I only found the root cause after working backward from the spike in false positives, validating field mappings one by one until I identified the upstream source. The pipeline itself wasn't broken. The data flowing through it was. That experience reinforced an important lesson: even the best detection models are only as reliable as the data they receive. Therefore it is safe to say that this is not a theoretical problem. It is a systemic vulnerability across financial crime compliance that a vast number of institutions encounter as systems and data evolve.
The industry is seeing this problem more often than many people realize. According to ComplyAdvantage's State of Financial Crime survey, 45% of compliance professionals name poor-quality, hidden data as the top barrier to effective risk detection. Separate insights from McKinsey & Company point to fragmented, legacy data as the major contributor behind the massive volume of false positives eating up investigators’ time.
I see firsthand how teams often spend weeks tweaking detection logic, only to discover later that the data itself was broken from the start. If we don't audit our incoming data as strictly as we write our logic the whole system becomes susceptible to failure.
Where AML Data Pipelines Break Down
In standard business intelligence pipelines, bad data might simply result in a skewed chart or minor metric discrepancy. However in an AML pipeline, it may actively prevent the software from catching compliance issues which may lead to regulatory penalties and legal issues.
Some of the most dangerous financial crime gaps stem from four persistent data failures:
1. Undetected Data Drift & Schema Mutations
One of the most difficult issues to detect is data drift. In my experience, engineers building an app or website are constantly pushing updates to support new business requirements. They may rename a field, change its format, or start sending null values instead of valid data.
What follows is this problem called Schema Mutation. The app team made their changes, started receiving weird data but just passed it along. The pipeline didn't break or fail, it just started writing garbage data into that column.
Data Drift is another common consequence: The structure is the same, but the actual data inside shifted. For example, a country field used to store full names (United States) but now stores codes (US). The logic (filter or lookup) on this field may miss this since the expected data was different.
Since the pipeline can still process the records without errors, these changes often go unnoticed.
2. Structural Mismatches and Fragmented KYC
Old, disconnected systems make it hard for banks to share data. Records for transactions, Know Your Customer (KYC) details, and risk profiles are housed in separate systems, often with inconsistent structures.
It can be hard to describe how bad the structure can get. In one system, all dates may be stored as alphanumeric strings, and in another, floating point numbers are used for Customer IDs, with vital AML fields missing. There may also be a case of a completely different risk score with a duplicate profile. Given how fragmented the data is, resolving profiles (the ability to tell whether or not "John Doe" and "J. Doe" are different people) is not even possible. In the end, this results in split profiles that conceal how a customer transacts, and causes other systems to miss suspicious behavior or generate an overwhelming amount of false alerts.
3. Structural and Semantic Blind Spots
A technically perfect system can still fail if the semantic meaning of the data changes. Structuring detection logic requires rigorous time-window joins—such as checking whether a customer conducted multiple cash deposits across different branches within a rolling 24-hour window to evade Currency Transaction Report (CTR) thresholds. If the timezone offset is dropped, truncated, or improperly cast during ingestion, the window shifts. The data remains clean and structurally sound, but the order of events is lost, making the review worthless.
4. The Upstream Disconnect: A Lack of Contractual Ownership
The most dangerous systemic risk is that AML pipelines heavily depend on upstream data sources owned by teams who have no idea their technical changes carry massive regulatory consequences. Because there is traditionally no shared ownership or mandatory validation between core banking engineers and compliance engineers, data integrity is broken before it even reaches the analytics layer. Pipelines quietly depend on upstream data sources that are completely unstable, undocumented, or modified without alert-impact simulations.
Building a Culture of Evidence-Based Compliance
Over the years, AML programs have focused heavily on creating detection rules and meeting regulatory requirements. While those are important, I've found that strong compliance starts with something much simpler—trustworthy data.
As data engineers, we spend a lot of time making sure information moves correctly from source systems to downstream applications. If that data is incomplete, inaccurate, or changes unexpectedly, even the best detection rules can produce unreliable results. That's why data governance and data quality should be part of every AML program, not an afterthought.
The industry is also moving toward a more evidence-based approach to compliance. Instead of relying only on predefined rules or thresholds, organizations are expected to demonstrate that their monitoring systems are built on accurate, complete, and well-governed data. A rule is only as effective as the data it evaluates.
Regulators are no longer interested in seeing only the alerts your system generated. They also want to understand why those alerts were generated—or why they were missed. That means you should be able to trace every important data element from the source system, through each transformation, to the final AML model or alert. Having that level of data lineage not only helps during audits but also makes it much easier to investigate issues when something doesn't look right.
Building More Reliable AML Pipelines: Engineering the Governance Framework
Fixing these vulnerabilities doesn't require a flashier detection rule or a more expensive black-box AI tool. It requires implementing precise, rigorous data governance and data engineering practices that treat data integrity as a hard regulatory requirement.
To safeguard an AML program against silent data corruption, institutions must build their technical strategy around these foundational pillars:
1. Programmatic Data Contracts and Ingestion Firewalls
To solve the upstream disconnect, organizations must implement Data Contracts. A data contract is a formal, version-controlled agreement between upstream software engineers and downstream data consumers that defines the expected schema, SLA, and semantic meaning of the data.
Coupled with this, automated Data Quality Firewalls must be pushed directly to the ingestion layer. Instead of cleaning data downstream in a warehouse or lakehouse, data must be validated the second it lands:
Schema Validation: Strict checking to ensure required AML fields (such as Originator/Beneficiary routing numbers or legal names) are populated and conform to explicit structures like SWIFT or ISO 20022 messaging standards.
Value Constraint Checking: Automated checks for null values, anomalous zero balances, or invalid data types (e.g., catching when a critical identifier accidentally converts to a float).
Quarantine Routing: Any data that violates the contract is automatically goes to a separate folder immediately so it cannot break the main system.
2. Immutable Storage and Deterministic Lineage Tracking
For an AML framework to be truly defensible, data cannot just be stored; it must be auditable across time. Using modern, open table formats that support ACID transactions and Time-Travel queries is essential.
Code:
[Upstream Core System]
│
│
(Data Contract)
▼
[Data Quality Firewall]
Check Schema & Constraints
/ \
(Pass) (Fail)
/ \
▼ ▼
[Curated Immutable [Quarantine Layer]
Tables] Alert Engineering
With Snapshot
Time-Travel
│
▼
[Deterministic Lineage Engine]
Traces Alert -> Code -> Source
Snapshot Isolation & Time-Travel: This capability allows engineers to query the data exactly as it existed at a specific millisecond in the past. If an investigator or an auditor asks why an alert is fired on a specific date, you can rewind data to see the exact moment of that customer's profile and transactions at that exact moment.
Deterministic Lineage: Every transformation—from raw string data to enriched customer profiles—must be mapped. Metadata platforms must capture the exact version of the detection code and the exact snapshot ID of the data used, ensuring absolute reproducibility.
3. Cross-Functional Data Discovery and Unit Testing
Before deploying or altering any critical compliance model, engineering and compliance teams must operate in lockstep through a disciplined validation process:
Deconstruct Intent: Technical teams must profile raw data sources and explicitly map fields to the regulatory basis of the rule (such as the Bank Secrecy Act, FinCEN guidance, or FINRA sales-practice rules) alongside the compliance analysts who actually work the alerts.
Edge-Case Unit Testing: Build modular transformations that are unit-tested specifically against dirty data: threshold boundaries, nulls, duplicates, and same-day reversals.
Historical Parallel Validation: When migrating systems or tuning thresholds, run old and new pipelines in parallel using a minimum of 90 days of historical data to compare outputs exactly.
Volume Guardrails: Establish baseline alert volumes. In my own deployments, I've used a drift threshold of roughly 20% against historical baselines as the trigger point for automated guardrails to halt a rollout—the exact number should be calibrated to each institution's own alert volatility, but the principle holds: production rollouts need an automatic circuit breaker before compliance operations are disrupted.
Conclusion: Trust Your Data Before Your Rules
One lesson I have learned is that an AML program is only as reliable as the data behind it. You can spend weeks refining detection rules and adjusting thresholds, but if the underlying data is incomplete or inaccurate, the results will never be reliable.
For me, improving data governance isn't just about building better pipelines; it's about building confidence in the decisions those pipelines support. Validating data, tracking lineage, and monitoring for unexpected changes help ensure that AML models continue to perform as intended.
Before changing another rule, ask yourself a simple question: Can I trust the data feeding it? If the answer isn't a clear yes, that's where the investigation should begin.
As AML systems become more data-driven, I believe data quality will play an even bigger role in successful compliance programs. Strong detection starts with trustworthy data, and that's something every data engineering team can help deliver.