Access to high-quality financial data has long been one of the biggest obstacles in improving anti-money laundering detection systems. By launching a synthetic dataset built from real banking patterns, the Financial Conduct Authority and the Alan Turing Institute are attempting to solve this challenge, enabling banks, fintech firms and technology vendors to train and test financial crime detection models without exposing sensitive customer information.
The Financial Conduct Authority (FCA) has officially released a statistically realistic synthetic dataset aimed at revolutionising how the financial sector detects and prevents money laundering. Developed in collaboration with the Alan Turing Institute, Plenitude, and Napier AI, the project addresses the historical challenge of accessing high-quality transaction data without compromising customer privacy. The team used anonymised real-world retail banking data from the United Kingdom to generate an entirely synthetic collection of customer profiles and financial transactions that mirror the statistical properties of genuine banking activity.
This initiative is a centrepiece of the FCA’s broader strategic goal to become a more data-led regulator and to reduce social harm caused by illicit financial flows. The project team employed advanced data generation techniques, including Generative Adversarial Networks (GANs) and differential privacy controls, to ensure that no individual or specific transaction can be reverse-engineered from the final product. By introducing controlled randomness, the dataset preserves the complex relational patterns essential for detection testing while strictly adhering to data protection laws and privacy standards.
The resulting dataset deliberately incorporates several realistic money laundering typologies observed across the financial sector. These include structuring payments to stay below reporting thresholds, rapid layering of funds across linked accounts, circular round-tripping, and high-risk cross-border transfers. To ensure the tools remain effective against evolving criminal tactics, the researchers introduced variations in these patterns to reflect the diversity and complexity of real financial crime. This approach shifts testing away from static rules toward models that can identify sophisticated behavioural ripples within the financial ecosystem.
Financial institutions and technology vendors can access this resource through the Digital Sandbox as part of the Synthetic Data AML Solution Sprint. This cohort-based program, running from May to July 2026, invites participating firms to demonstrate how emerging technologies, such as artificial intelligence and machine learning, can enhance the effectiveness of transaction monitoring. The sprint provides a safe, standardised environment for benchmarking new solutions, allowing innovators to validate their tools without needing sensitive live banking records.
Napier AI, a key partner in the project, recently showcased how these types of datasets were used to develop its new Insights AI feature, which uses frequency-based algorithms to detect suspicious flows. Dr Janet Bastiman, Chief Data Scientist at Napier AI and a member of the FCA Synthetic Data Expert Group, noted that the project overcomes the disconnected nature of data required for pattern analysis throughout the lifecycle of a transaction. As the deadline for sprint applications approaches on April 26, the FCA expects the findings to provide critical evidence on how synthetic data can be used in regulatory supervision and the wider fight against economic crime.
What this means for the industry
- Synthetic data could transform AML innovation
Financial institutions have historically struggled to test detection models because real transaction data is restricted by privacy regulations. Synthetic datasets allow developers to build and refine AML tools while maintaining strict data protection standards. - Regulators are becoming technology partners
The FCA’s involvement signals a shift toward regulators actively supporting innovation rather than simply enforcing compliance. By providing testing environments like the Digital Sandbox, regulators are encouraging collaboration with fintech and AI developers. - AI-driven financial crime detection is accelerating
Techniques such as machine learning and behavioural analytics require large, complex datasets. Synthetic data enables firms to train these models to detect sophisticated laundering patterns such as layering, structuring and cross-border flows. - Standardised datasets enable better benchmarking
When institutions test their systems using a common dataset, it becomes easier to compare detection performance across different technologies, improving the industry’s overall ability to combat financial crime. - Privacy-preserving analytics is becoming a regulatory priority
The use of technologies such as Generative Adversarial Networks and differential privacy reflects a growing focus on balancing financial surveillance with strict data protection requirements.

