When The New York Times sued OpenAI and Microsoft in late 2023, the publisher aimed straight for the jugular: billions of dollars in statutory damages and the court-ordered destruction of foundational models trained on its archives. It was framed as an existential showdown over intellectual property. Now, the US Department of Justice has formally weighed in—and delivered a decisive blow to the media industry's monetization strategy.
The Training-versus-Output Distinction
In its amicus brief, the DOJ explicitly drew a line between computational data ingestion and infringing generation. Siding with model developers, federal attorneys argued that ingesting copyrighted text for model training falls squarely under the fair use doctrine. The government pointed out the core flaw in the publishers' legal theory: conflating the internal process of algorithmic training with the final output delivered to enterprise users.
While models process full source texts to extract statistical patterns and semantic relationships, the underlying works are not distributed to the public. As long as the resulting generations lack substantial similarity to the source data, the training phase itself does not constitute actionable infringement. To sharpen the point, the DOJ cited Joan Didion, who famously typed out Hemingway's prose word-for-word as a teenager to master sentence structure. Under the Times' theory, Didion would have incurred liability the moment she published her own novels. Imposing copyright penalties on algorithmic learning, the government warned, would systematically throttle the innovation that copyright law was designed to incentivize.
Administrative Clashes Over AI Training Scale
The DOJ's intervention puts it in sharp conflict with the US Copyright Office, which previously argued that a multibillion-dollar foundation model cannot be compared to an individual author studying a text. Yet for enterprise buyers and ML leaders, the DOJ's stance provides crucial legal clarity.
Mandating retroactive licensing deals for every token in an open-web scrape would create a catastrophic financial bottleneck, effectively reserving advanced LLM development for a handful of mega-caps capable of paying off legacy media conglomerates. By defending the fair use status of model ingestion, the DOJ significantly lowers the risk of structural liability and prohibitive dataset royalties across parallel copyright dockets. The debate will shift where it belongs: from whether models can learn from public data to whether their commercial outputs unlawfully copy protected expressions.