Autonomous language-model agents are actively deployed with real capital across live financial order flows, yet their empirical behavior under market pressure has remained largely unmeasured at scale. DX Research Group analyzed a six-month continuous production record across two distinct agent fleets sharing a common software architecture. The underlying dataset encompasses 7.5 million single-model invocations, roughly 300,000 onchain actions, and 231,638 multi-tool turns that yielded just 14,596 executed fills across Base memecoin markets and Hyperliquid perpetuals.
The Realities of Production Fleets
The evaluated architectures captured two separate live operational phases during 2026. From February to March 2026, the DX Terminal Pro system ran 3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days. Subsequently, from June to August 2026, the DXAP live alpha fleet operated between 500 and 599 user-created agent instances trading Hyperliquid perpetual contracts.
"Neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark"
As the DX Research Group team reported in their paper, the DXAP fleet generated no structural alpha, delivering a meager 41% roundtrip win rate and trailing behind a matched Hyperliquid retail benchmark that posted a 50% win rate. Furthermore, evaluating frontier models across 416 captured production scenarios in a paired-replay league revealed that decision quality remained statistically indistinguishable across leading model families at that execution horizon.
Structural Constraints and Sizing Blindness
Across both deployment fleets, quantitative regression demonstrated that the operating environment and deterministic parameters governed execution patterns far more strongly than prompt engineering or strategy text. In the empirical data, the system risk slider accounted for leverage scaling at a rate of +0.425 per increment level, and a leaderboard render boundary causally drove capital allocation with a 1.75× regression discontinuity at the top-3 cutoff threshold.
Crucially, agent sizing logic remained blind to shifting market volatility. Median leverage settled stubbornly at 5.0× across every single volatility sextile without dynamic adjustment. Consequently, extreme portfolio vulnerability concentrated within rigid risk configurations: one posture-slider cell representing 11% of the total book accounted for 62% of all recorded liquidations.
Execution Capture and Practical Significance
The autonomous decision loops struggled heavily with trade lifecycle execution and exit discipline. While 43.2% of open agent positions achieved a favorable excursion of at least +300 basis points within a 24-hour window, 49.3% of those winning positions ultimately closed with an overall negative trade return. The models consistently failed to secure unrealized gains before market reversal dynamics wiped out temporary profits.
For enterprise architects and fintech leaders, the takeaway is unequivocal: deploying autonomous LLMs into high-frequency, capital-allocating domains without rigid, deterministic execution guardrails is an invitation to capital destruction. Unbounded model reasoning cannot replace hardcoded risk rules, dynamic volatility sizing, and deterministic stop-outs.