Evaluating LLM agents on synthetic reasoning puzzles or isolated tool workflows creates a dangerous illusion of enterprise competence. In real business operations, capital allocation requires orchestrating pricing, production, procurement, inventory, and treasury loops where early miscalculations compound into financial insolvency. To quantify how autonomous decision-makers actually handle shared, competitive market dynamics, researchers from Shanghai Innovation Institute, Beijing Institute of Technology, Shanghai Jiao Tong University, and GAIR Lab introduced ERPBench—an execution-instrumented stress test for enterprise resource planning.
Multi-Round Compounding and Market Structure
ERPBench tracks agent performance across a six-round enterprise simulation governing five coupled operational domains: dynamic pricing, procurement schedules, production capacity, inventory management, and treasury balance. Each decision cycle compounds directly into the next, ensuring that early missteps in inventory overstocking or price undercutting ripple through the company's financial runway.
In competitive markets, strategic choices are tightly coupled: a single agent's aggressive pricing or inventory dump fundamentally alters the addressable demand for everyone else.
The benchmark measures ultimate enterprise viability through terminal company valuation across 100 standardized scenarios. Crucially, it splits testing across two distinct market environments: a "Solo" setting where an LLM agent faces predictable, rule-based opponents, and a cutthroat "Arena" mode where six LLM agents compete directly within the same shared-market economy.
Leaderboard Instability Across Competitive Environments
Covering six frontier model families across 1,200 trajectories and 7,200 operational decision rounds, the empirical data shatters the assumption that isolated performance benchmarks predict real-world competitiveness.
DeepSeek dominated the Solo environment with a 252.29M mean valuation and an average rank of 1.67, but stumbled when facing adaptive peers. Conversely, Gemini took the top spot in the Arena ecology, posting a 263.95M mean valuation and a 1.76 mean rank, while dropping its bottom-rank failure rate from 22% in Solo to zero in Arena. Across the board, Solo rankings matched Arena outcomes on only 21 out of 100 evaluated scenarios.
Static, single-agent benchmarks fail to reflect the cascading operational failures that occur when autonomous models encounter competitive feedback loops. Deploying autonomous agents into enterprise supply chains and treasury operations without multi-agent dynamic stress-testing is a direct path to unhedged balance sheet destruction.