The price of hitting fixed performance thresholds across standardized artificial intelligence benchmarks has plummeted since 2023. Tracking groups monitoring frontier model pricing observe a relentless collapse in costs, though the true magnitude depends heavily on whether analysts filter out market noise from underlying engineering gains.
Market Drops Versus Pure Efficiency
Research organization Epoch AI, tracking trends across five benchmarks spanning mathematics, science, and logic puzzles, calculates that inference costs are dropping by roughly 47 percent per quarter—a staggering 13-fold deflation rate annually. While these figures represent rough approximations based on available market data, no previous industrial technology has ever scaled cost reductions at this velocity.
Epoch AI calculates that inference costs are dropping by about 47 percent per quarter on average, or roughly 13x per year.
Yet this headline rate reflects aggressive market pricing and hardware subsidies rather than pure algorithmic breakthroughs. When independent researchers, including Hans Gundlach and his colleagues, evaluate pricing data from the comparison platform Artificial Analysis spanning April 2024 through November 2025, total cost reductions land closer to 5x to 10x annually.
Isolating Algorithmic Gains
The gap between these metrics comes down to external variables. Epoch's 13x figure runs higher because it fails to strip out the compounding impact of cheaper silicon and brutal price wars among infrastructure vendors. Once analysts control for hardware efficiencies and market competition, the underlying gain in algorithmic efficiency stabilizes at roughly 3x per year.
At the same time, benchmark progress obscures a critical strategic fork in the road for engineering leaders. While running legacy model tiers grows cheaper by the month, advanced reasoning architectures consume substantially more compute per execution. For instance, OpenAI's o3 model achieved a 75 percent score on the GPQA Diamond PhD-level science test at an estimated 30 cents per query. Matching historical performance levels gets cheaper over time, but deploying top-tier reasoning models drives compute budgets through the ceiling.
Headline deflation figures mask the real mechanics at play: market competition and hardware scale drive most of the near-term price drop, while true algorithmic efficiency grinds forward at a steady 3x per year, forcing technical leads to balance cheap commodity queries against expensive, high-stakes reasoning pipelines.