Allocating extra test-time compute has become a standard engineering reflex for squeezing better outputs out of large language models on benchmarks with immediate feedback, such as coding and mathematics. Applying internal reasoning budgets to noisy financial markets, however, presents a harsher reality where success is measured strictly after real-world execution costs. In a preprint titled "The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?", Jiayi Chen and Guiling Wang from the Department of Computer Science at the New Jersey Institute of Technology evaluated whether increasing reasoning effort inside an otherwise fixed trading pipeline actually produces superior portfolios.

Controlled Backtesting Framework

To isolate the direct economic impact of extra inference tokens, the researchers locked down information availability, prompts, required output formats, and portfolio construction rules across evaluation runs. The study tracked representative models from three predetermined families—DeepSeek, GPT, and Gemini. During each trading day, each model ranked the same equities under three separate input conditions: numerical data, identifiable news, and masked news.

"Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns."

In total, the benchmark processed more than 800,000 asset predictions alongside repeated generations to test for stochastic variance. The experimental design established baseline reasoning as zero for DeepSeek and GPT, while setting Gemini to minimal reasoning because its API architecture cannot guarantee that internal thinking loops are completely disabled.

Nonmonotonic Gains and Selection Instability

To put it bluntly, throwing more compute at the problem does not translate into proportional financial gains. Across all three model families, additional reasoning failed to deliver a reliable bump in net portfolio returns after factoring in real transaction costs. For the DeepSeek model, where the authors evaluated the complete trajectory from no reasoning to maximum reasoning, performance proved frustratingly nonmonotonic rather than steadily improving.

Furthermore, repeated generations produced unstable treatment effects and scrambled portfolio asset selections even when overall model scores appeared consistent. While test-time reasoning visibly altered internal stock scores, failure patterns, and operational token bills, those theoretical shifts routinely dissolved the moment execution expenses hit the books.

This looks like a sobering reality check for quantitative funds. The findings establish that spending additional compute at inference time changes decision trajectories without guaranteeing measurable economic value in noisy domains. Because trading performance degrades across stochastic generations and transaction friction, engineering teams cannot simply assume that benchmark reasoning gains will transfer automatically to live financial pipelines. Validating test-time reasoning on domain-specific execution metrics remains an absolute requirement before committing serious compute budgets to production.

Artificial IntelligenceLarge Language ModelsAI in FinanceOpenAI