The financial sector is grappling with a paradox: the inherent flexibility of Large Language Models (LLMs) is clashing head-on with rigid regulatory requirements. While LLM-powered chatbots attempt to navigate complex customer queries, banks are still stuck in the past, testing them via manual focus groups. This traditional approach is slow, expensive, and fundamentally incompatible with iterative development. Researchers at NatWest AI Research, led by Cristóvão Iglesias, have proposed a way out: the concept of Synthetic Customer Agents (SCA)—digital twins designed to replace the army of human testers.
Foundations in Real Transactions: How Twins Are Built
The methodology presented by the NatWest team at the ICLR 2026 workshop goes far beyond simple persona creation via ChatGPT. It is a two-component framework: first, the system digests dialogue histories, request contexts, and real transactional data to transform them into structured system prompts. This ensures the synthetic agent doesn't just mimic human speech but acts based on a specific customer’s financial background. As the researchers note, this allows for the conditioning of AI behavior, simulating a wide range of profiles and communication styles.
"Synthetic agents achieve high semantic alignment with real customers, demonstrate low hallucination rates, and successfully reproduce personality traits under controlled interventions."
Using transcripts of real conversations allows for the recreation of typical behavior. However, to stress-test a bot’s resilience, the researchers apply personality conditioning. This allows them to dial up specific emotions in the digital twin—making it aggressive, confused, or excessively demanding. Consequently, the bank can run its chatbot through multi-turn dialogues where the "customer" is intentionally difficult, verifying whether the AI remains compliant under pressure.
Automated Validation and Regulatory Oversight
The second pillar of NatWest’s work is a scalable evaluation system where one LLM acts as a judge for another, utilizing adversarial attacks in the process. By putting a chatbot through thousands of these synthetic interactions, the bank can identify edge cases and hallucinations that a human reviewer might miss due to fatigue or oversight. According to NatWest AI Research, this method has already been trialed on a customer-facing bot at one of the UK's largest banks. For regulators like the FCA (Financial Conduct Authority), this framework provides a compelling argument: it assesses AI performance across all demographics and emotional states, ensuring an absence of bias.