For two years, the foundation model market endured well-founded skepticism over its structural burn rate. Early unit economics saw frontier labs torching billions on compute without demonstrable positive contribution margins on token delivery. That narrative is changing as model providers reorganize around industrial fundamentals: purchasing wholesale electricity at $10 million to $15 million per megawatt and packaging it into monetizable cognitive work.
The Turnaround in Megawatt Economics
The real cost floor in AI infrastructure begins with physical power and data center capacity. Dylan Patel of SemiAnalysis noted on the Dwarkesh Podcast that baseline compute operations cost between $10 million and $15 million per megawatt. In 2024, Anthropic operated with an unsustainable gross margin of −94%, burning $1.94 in infrastructure for every single dollar collected.
That margin equation inverted as enterprise inference demand matured and operational execution improved. Anthropic expanded its gross margin into positive territory, generating up to $50 million in top-line revenue per megawatt against its $10 million to $15 million cost foundation. That shift enabled Anthropic to book $559 million in quarterly operating profit on $10.9 billion in annualized run-rate revenue, turning inference capacity into a self-sustaining cash engine.
"The base cost of compute tends to be around 10 or 13 or $15 million per megawatt. In the case of Anthropic, the revenue has gone as high as $50 million per megawatt. And what that now enables them to do is, hey, if I spend 10 bucks on inference capacity, I actually generate 50 bucks of revenue. And then I can turn around and incrementally spend all of that profit on training."
This dynamic restructures frontier AI financing. While Anthropic's net operating margin of ~5% sits well beneath theoretical 80% SaaS gross margins, the gap reflects aggressive reinvestment rather than structural leakage. Operating profits generated on token sales are funneled directly into pre-training clusters and engineering headcount, establishing an internal capital flywheel that reduces dependence on venture tranches.
Architectural Efficiency and the Model Factory
Squeezing $50 million from a megawatt is fundamentally an engineering problem. Sparse architectures and optimized inference kernels allow providers to deliver target intelligence benchmarks at an order-of-magnitude reduction in active parameter load, driving down Joules-per-token and protecting API pricing stability for downstream enterprise buyers.
Model development has rapidly matured from speculative research into an automated industrial pipeline. Nvidia's $6 billion acquisition of Poolside—backed by a further $1 billion operational commitment—reflects this structural transition. As Poolside's Jason Warner pointed out, individual frontier models are ephemeral outputs; the durable enterprise asset is the underlying model factory capable of automating architecture search and shipping production weights like Laguna S in weeks rather than quarters.
Inference is no longer a subsidized loss-leader to win market share. For business leaders building on foundational APIs, high megawatt yields signal reliable pricing floors and ensure that the vendors providing their cognitive infrastructure can fund next-generation model training entirely out of operating cash flow.