For decades, enterprise compute followed an inevitable trajectory away from centralized mainframes toward distributed silicon. Koomey’s law captured that structural shift by tracking how computing power per watt doubled roughly every 1.5 years, eventually condensing industrial-grade workloads onto desktop hardware. Enterprise AI is now reaching its own decentralization inflection point, with inference shifting steadily away from hyperscaler data centers and back behind corporate firewalls.

According to Tomasz Tunguz, evaluating benchmark research by Jon Saad-Falcon, Avanika Narayan, and their co-authors from Stanford University and Together AI, local deployments already match frontier cloud systems across 89% of standard chat and reasoning workloads. Based on an evaluation across more than one million real-world queries and 20-plus local models, the overall efficiency of running smaller models locally has increased 5.3x in just two years.

Efficiency Gains and Model Routing

Hardware improvements explain only part of this economic shift. As Anson Ho, Ege Erdil, and Tamay Besiroglu demonstrated in their 2023 research on CMOS scaling limits, GPUs doubled their computational efficiency roughly every 2.7 years over the past decade and a half. The recent 5.3x jump in intelligence-per-watt combines a 1.7x efficiency dividend from denser silicon with a 3.1x leap from leaner model architectures.

Between 2023 and 2025, the win-or-tie rate of the top standalone local model against frontier cloud endpoints jumped from 23.2% to 71.3%, compounding at roughly 20 percentage points annually. Pairing compact open weights with an automated local router pushes that frontier higher by dynamically assigning each prompt to the most efficient specialized model.

"Just as performance-per-watt guided the mainframe-to-PC transition, intelligence-per-watt will guide AI’s transition to the edge."

As the Stanford and Together AI researchers observed, local routing across multiple specialized engines lifts the win-or-tie ceiling to nearly 90%, effectively commoditizing routine enterprise queries.

Infrastructure Trade-Offs

Centralized server farms retain a decisive moat for edge-case reasoning, deep domain synthesis, and massive parallel agentic execution. In pure hardware utilization, hyperscalers still capture a 40% energy efficiency edge over local silicon thanks to continuous query batching at massive scale.

Yet for standard enterprise workflows, sending everyday knowledge work over external APIs is becoming an unjustifiable operational tax. Keeping inference on-premise eliminates off-site latency, ring-fences sensitive enterprise IP, and fundamentally resets TCO. Across standard operational workloads, a local-plus-router architecture cuts energy consumption by 80%, slashes compute overhead by 77%, and reduces total infrastructure spend by 74% compared to an all-cloud baseline.

On-Device AICost ReductionAI in BusinessOpen Source AICloud Computing