The AI market is witnessing a tectonic shift toward cheap, commoditized reasoning. According to testing results dated July 31, 2026, the DeepSeek V4 Flash 0731 model has delivered a standout performance on the ARC-AGI benchmark—the gold standard for measuring an AI’s ability to generalize across unfamiliar tasks. By hitting 89.0% on the semi-private ARC-AGI-1 set and 61.4% on ARC-AGI-2 at 'Max effort,' DeepSeek suggests its architecture can actually solve novel problems rather than just regurgitating training patterns.

Three-Tier Reasoning Architecture

The methodology behind V4 Flash 0731 relies on three compute-intensity settings: Max, High, and Low. The data reveals a brutal correlation between compute spend and accuracy. While the Max variant holds the line at 89.0% on ARC-AGI-1, dropping to Low sinks performance to 84.0%. On the more grueling ARC-AGI-2, the gap widens into a chasm: 61.4% for Max versus a mediocre 46.0% for Low.

At Max effort, DeepSeek V4 Flash 0731 hits 89.0% on ARC-AGI-1 Semi-Private for a mere $0.02 per task.

This granular control allows for surgical resource allocation based on task complexity—a prerequisite for real-time autonomous systems. Interestingly, public evaluations of ARC-AGI-2 show that while Max conquered tasks like 135a2760, it stumbled on 0934a4d8—a task the High variant actually passed. This highlights the non-linear nature of logical inference: throwing more raw power at a problem doesn't always guarantee a solution.

The Economics of Inference vs. Raw Power

DeepSeek’s real triumph isn't just the percentage of correct answers; it's the aggressive devaluation of the 'unit of intelligence.' Solving a task on ARC-AGI-1 costs $0.02, rising to only $0.04 for the demanding ARC-AGI-2. In the context of the François Chollet test, which targets abstract reasoning, this price point changes the math for deploying autonomous logic agents at scale.

The model demonstrates genuine generalization on semi-private datasets, resolving tasks at a cycle cost between $0.02 and $0.04.

By turning sophisticated reasoning into a cheap consumable, DeepSeek is making mass adoption of autonomous logic agents economically viable. While Western models of a comparable class often bleed cash on high-inference costs, the Chinese team is optimizing the cost of 'thought' itself. This moves the AGI conversation from theoretical benchmarks to operational bottom lines. The catch remains the reliance on semi-private datasets; the true test will be whether this 89% accuracy survives in the wild, industrial environments where there are no pre-calibrated safety nets or hidden data overlaps. For now, DeepSeek isn't just competing on power—it's winning on price.

Large Language ModelsAI in BusinessCost ReductionDeepSeek