Conversational banking agents fail between 35% and 51% of adaptive fraud attacks when granted operational tools alongside internal policies, according to a benchmark by researchers Dheeraj Mohandas Pai and Lu Xian from Leanmcp. Evaluating four enterprise agents across 107 graded adversarial scenarios revealed that baseline attack-security rates hover between an alarming 49% and 65%. Money-mule schemes and first-party fraud proved to be the most damaging failure modes across all evaluated architectures.

The findings originate from FraudBench, an evaluation environment built on the τ2-bench dual-control architecture and the τ-Knowledge banking ecosystem. To replicate standard customer service deployments, agents were armed with operational APIs to execute funds transfers, reset authentication PINs, and modify contact records, while backed by a 698-document internal policy repository. Standard enterprise testing focuses on static transaction analysis or brute-force prompt injections; FraudBench instead stresses multi-turn conversational manipulation where bad actors progressively game authorization boundaries across extended dialogues.

The systemic architectural flaw lies in coupling client-facing conversational interfaces with back-office execution tools. When support agents are given direct authority to process state-changing transactions, single-control workflows leave them hopelessly exposed to adaptive probing and chained requests that appear benign in isolation. Expecting conversational safety guardrails to double as a hard fraud desk is wishful thinking. Fintech leaders must enforce isolated out-of-band verification perimeters before handing autonomous models execution rights over enterprise financial pipelines.

AI AgentsAI in FinanceAI SafetyCybersecurity