The economics of training large language models has long shifted from a battle of giant compute clusters to a basic shortage of quality text corpora. The public internet has been scraped dry, and developers are now forced to buy closed archives, corporate mailboxes, and specialized knowledge bases for millions of dollars. This fundamentally redefines the profitability of model fine-tuning and forces businesses to recalculate their infrastructure economics.
Paid Access to User Discussions
Forums have accumulated living human dialogues for decades, so model developers view them as an ideal source of texts. The first notable signal of changing rules was Reddit. In April 2023, the platform announced the cancellation of free access to its API to protect content from uncontrolled use by LLM creators. The most popular iOS client called Apollo shut down because its developer admitted the API would cost about $20 million a year.
Google's access to Reddit discussions cost roughly $60 million a year, demonstrating the real price of user text data for model developers.
Preparing for its IPO, Reddit management disclosed that the total volume of such agreements exceeded $200 million. The market has definitively transitioned to direct commercial purchases of text arrays.
Contracts with Publishers and Archives
In the spring of 2024, OpenAI began systematically signing licensing contracts with media outlets and archive holders. The company finalized an agreement with News Corp worth over $250 million for a five-year term. This was followed by contracts with media holding Dotdash Meredith for at least $16 million, Axel Springer for roughly $13 million annually for up to three years, and the Financial Times for $5–10 million annually.
Similar processes affected scientific literature and visual content. Wiley executed two deals with unnamed partners worth $21 million and $23 million respectively, while Britain's Informa received an initial payment of $10 million from Microsoft.
Open Internet Exhaustion and Bankruptcy Auctions
Massive spending is driven by the demographics of training datasets. According to a 2024 report by researchers at Epoch AI, human-created public content will begin to run out sharply by 2026. This shortage of quality raw material forces corporations to look for data in the most unusual places, including closed corporate archives.
In our view, buying other people's correspondence and knowledge bases will inevitably run into severe legal risks, compliance hurdles, and privacy protection issues. The sale of corporate archives can flip the entire economics of model fine-tuning, turning legal departments into the primary regulators of the AI market. Corporations will have to choose between lawsuits for data leaks and billion-dollar investments in closed datasets.