top of page

A Sandbox Result Is Not a Supervisory Blessing: Supervised Testing of Financial Crime Technology

Writer: TrustSphere Network
TrustSphere Network
11 minutes ago
8 min read

The supervisory posture towards financial crime innovation has settled into an awkward but stable position. Supervisors want firms to adopt better technology and say so publicly. They are simultaneously unwilling to pre-approve tools, unwilling to be seen to endorse a vendor, and acutely aware that a poorly governed model deployed against customers produces harms they will be asked to answer for. The instrument that has emerged to hold both positions at once is supervised testing: environments where a firm can try something under observation, with conditions attached, without the consequences of production.


For financial crime that instrument has become more relevant than it was. The controls now being proposed are probabilistic rather than deterministic, frequently built on general purpose models the firm does not control, and dependent on data the firm may not lawfully hold in the volume required to test them properly. A tuning change to a rules engine can be validated on historical data in a controlled way. A language model performing enhanced due diligence, or a network analytic operating across data contributed by several institutions, cannot be, and the alternative of learning in production is exactly what a supervisor does not want.


What follows sets out what actually exists, what is being consulted on and what is speculation, and is then direct about the operational consequences. The most important of those is also the most frequently misunderstood: a sandbox outcome is a test result. It is not a permission, not a waiver, and not a supervisory opinion that the firm's control is adequate. Institutions that blur that distinction create a governance exposure worse than the one they were trying to resolve.


What Is Changing


Begin with what is established. The Financial Conduct Authority has operated a regulatory sandbox for around a decade, admitting firms to test propositions with real customers under restrictions and with a dedicated supervisory contact. Alongside it sits a digital sandbox built around synthetic data assets, a permanent service rather than a pilot, aimed at firms that need data to prove a model before they have the customers or permissions to generate it themselves. None of this is new, and financial crime use cases have featured for years, particularly around identity, onboarding and scam detection.


What has shifted is the emphasis, in three directions. The first is synthetic and shared data. The binding constraint on financial crime model development is not algorithms, it is representative labelled data, and confirmed financial crime outcomes are rare, sensitive and unevenly distributed across institutions. Supervisory and public sector interest in synthetic datasets and shared testbeds follows directly from that, and the international central banking community has pursued the same line: the Bank for International Settlements Innovation Hub has run projects examining whether privacy-preserving analytics across data from multiple institutions can detect money laundering patterns that single-institution monitoring misses. Those projects report findings rather than confer permissions, and should be read as evidence about feasibility, not as regulatory change.


The second direction is information sharing. In the United Kingdom the Economic Crime and Corporate Transparency Act created direct gateways permitting firms in scope to share customer information with each other for economic crime purposes, subject to conditions. That is legislation in force rather than a pilot, and it is the foundation on which more ambitious collaborative analytics could be built. It does not by itself authorise pooling data into a shared model, and the gap between a permitted disclosure between two firms and a lawful multi-institution analytical environment is precisely where supervised testing has become useful. Firms have explored that gap through pilots with public sector participation, and supervisors have been more willing to observe such work than to bless its output.


The third direction is artificial intelligence. Supervisors have signalled consistently and in public a preference for observing AI-enabled tooling under controlled conditions before it reaches customers at scale, positioning testing services and innovation work accordingly rather than issuing prescriptive AI rules for financial services. In the European Union the position is codified: the Artificial Intelligence Act requires member states to establish national AI regulatory sandboxes and sets out how they operate, including testing under conditions and the documenting of participation. That is a legal obligation on member states rather than a discretionary supervisory initiative, and it is the most significant structural difference between the European and United Kingdom approaches. Other jurisdictions operate innovation facilities of their own, but the maturity and legal effect of participation vary considerably.


Timelines and What Is Still Uncertain


Treat the following as settled. Regulatory and digital sandbox facilities exist in the United Kingdom and are open on a standing basis. The Economic Crime and Corporate Transparency Act information sharing provisions are law. The European Union's AI Act contains binding requirements on member states to establish AI regulatory sandboxes, and its obligations for high-risk systems bring formal expectations on risk management, data governance, logging, human oversight and post-market monitoring that apply whether or not a firm enters a sandbox. The Prudential Regulation Authority's principles for model risk management apply to models in scope regardless of where they were developed or tested. None of these depends on a future decision.


What is uncertain divides into two categories worth keeping apart. The first is design detail genuinely being worked out: the scope and terms of shared and synthetic data assets, the conditions attached to multi-firm analytical pilots, how AI sandbox participation in member states will interact with high-risk conformity obligations in practice, and how consistently national authorities will run the facilities the AI Act requires. The second category is speculation, and we label it as such. We see no adopted instrument in any major jurisdiction that converts sandbox participation into a defence, a safe harbour or a supervisory approval for a financial crime control, and no published proposal to create one. The frequently voiced expectation that a certification pathway for AI-enabled financial crime tooling will emerge is, at present, an expectation and nothing more. Plan on the assumption that testing informs a firm's own decision and does not transfer responsibility for it.


What It Means Operationally


The first consequence is data protection and legal basis, and it will consume more time than the technical work. Testing a financial crime model on real customer data requires a lawful basis and a purpose compatible with the one for which the data was collected, documented before testing starts rather than reconstructed afterwards. Synthetic data is the obvious route around this, but synthetic is not automatically anonymous: a dataset generated from real records can retain re-identification risk, and the Information Commissioner's Office position on anonymisation and identifiability applies to synthetic outputs as to any other data. Shared environments raise harder questions again: controllership, the legal basis for each disclosure, retention and deletion at test end, and international transfer. A data protection impact assessment is not a formality here; it is the artefact that decides whether the test can lawfully proceed.


The second consequence is model governance evidence, and the useful reframing is that a supervised test is an opportunity to generate it. A model risk function will ask for an intended use statement, a validation approach, a performance measurement design, a monitoring plan and a statement of limitations. A well-constructed sandbox exercise produces all of those as by-products; one designed as a demonstration produces none of them. That means specifying in advance the success criteria, the control comparison, how ground truth is established, what constitutes a failure, and what evidence is retained. Synthetic data introduces a further obligation firms routinely skip: an explicit assessment of how well the synthetic distribution represents the production population, because a model validated against data that under-represents the rare cases it exists to catch has been validated against the wrong problem.


The third consequence is understanding what a test result legally is. A sandbox outcome is evidence the firm may rely on in its own decision making. It is not individual guidance, not a waiver or modification of a rule, not authorisation for an activity, and not a supervisory opinion that a control is adequate. Where a firm needs certainty about how a rule applies to it, the mechanism is the formal route for individual guidance from its regulator, and the answer is narrow, conditional and specific to the facts presented. Be equally clear internally about what supervisory engagement during a test means: an observing supervisor who raises no objection has not approved anything, and a board pack should say so.


The fourth consequence is the governance risk created by the language people use afterwards. The failure mode is predictable: a product is described internally as regulator-tested, the description migrates into a business case, then a board paper, then a control owner's justification for reduced assurance, and eventually a customer-facing claim. At each step the qualification degrades, and the institution ends up with a control whose adequacy rests on a test conducted under conditions that no longer hold, against a data distribution that has moved. The mitigations are unexciting and effective: a written record of what was tested and under what conditions, an explicit statement of what the test did not cover, a named owner for the re-validation that follows any material change, and a standing prohibition on describing sandbox participation as approval.


Conclusion


Supervised testing is a genuine and underused route for financial crime innovation, particularly for institutions whose second line will not accept an unvalidated probabilistic control and whose data holdings are too thin to validate it alone. It offers a structured way to generate evidence, a reason to define success criteria before building, and access to data assets a single firm cannot assemble.


Firms that approach it as a route to regulatory cover get the opposite. The settlement is deliberate and stable: the regulator provides the environment, the observation and sometimes the data, and the firm retains every element of responsibility for the control it subsequently deploys. That bargain is reasonable, but it only works if the institution is disciplined about what the test proved and equally disciplined about saying so afterwards. The most valuable output of a sandbox is usually not the model. It is the written record of the conditions under which the model was shown to work, and the honest list of the conditions under which it was not.


Suggested Next Steps


  • Complete the data protection analysis before any test design work, covering lawful basis, purpose compatibility, controllership in shared environments, retention and deletion at test end, international transfer, and an explicit assessment of re-identification risk in any synthetic dataset.


  • Design every supervised test to produce model governance artefacts as by-products, specifying success criteria, the control comparison, how ground truth is established and what evidence is retained, and agree the design with model risk management and internal audit before entry rather than after exit.


  • Assess and document how well any synthetic or shared dataset represents your production population, with particular attention to the rare positive cases the model exists to detect, and record the resulting limitation in the model's statement of intended use.


  • Adopt a written rule that sandbox or pilot participation is never described as regulatory approval in any internal or external document, appoint a named owner for re-validation following material change to the model, the data or the environment, and use the formal individual guidance route where you need certainty on how a rule applies.


Sources: Financial Conduct Authority regulatory sandbox, digital sandbox and innovation services, and publications on the use of artificial intelligence in financial services; Financial Conduct Authority financial crime guide and expectations on the effectiveness of systems and controls; Prudential Regulation Authority principles for model risk management; Economic Crime and Corporate Transparency Act information sharing provisions; His Majesty's Treasury economic crime plan and associated policy statements; Information Commissioner's Office guidance on anonymisation, pseudonymisation, data protection impact assessments and artificial intelligence; European Union Artificial Intelligence Act provisions on regulatory sandboxes and high-risk system obligations; European Banking Authority guidance on money laundering and terrorist financing risk factors; European Union anti-money laundering package and the establishment of the Anti-Money Laundering Authority; Bank for International Settlements Innovation Hub projects on privacy-preserving analytics and anti-money laundering; Financial Action Task Force work on digital transformation and the risk based approach; Wolfsberg Group statements on effectiveness and on the use of technology in financial crime programmes; TrustSphere Risk Index, April 2026.


TrustSphere helps financial institutions design and deploy intelligent fraud and financial crime detection solutions. Visit www.trustsphere.ai


 
 
 

Comments


Recommended by TrustSphere

© 2026 TrustSphere.ai. All Rights Reserved.

  • LinkedIn

Disclaimer for TRUSTSPHERE.AI

The content provided on the TRUSTSPHEREAI website is intended for informational purposes only. While we strive to provide accurate and up-to-date information, the data and insights presented are generated from a contributory network and consolidated largely through artificial intelligence. As such, the information may not be comprehensive, and we do not guarantee the accuracy, reliability, or completeness of any content.  Users are advised that important decisions should not be made based solely on the information provided on this website. We encourage users to seek professional advice and conduct their own research prior to making any significant decisions.  TruststSphere Partners is a consulting business. For a comprehensive review, analysis, or support on Technology Assessment, Strategy, or go-to-market strategies, please contact us to discuss a customized engagement project.   TRUSTSPHERE.AI, its affiliates, and contributors shall not be liable for any loss or damage arising from the use of or reliance on the information provided on this website. By using this site, you acknowledge and accept these terms.   If you have further questions,  require clarifications, or requests for removal or content or changes please feel free to reach out to us directly.  we can be reached at hello@trustsphere.ai

bottom of page