Destructive Scanning at Scale
The race to feed frontier language models has hit an inevitable wall on the open web, driving corporate data procurement teams back into the physical world. As publicly indexed internet text becomes contaminated with recursive synthetic scrapings, labs are paying a premium for pristine, pre-2022 human prose preserved on paper. To ingest physical volumes at high throughput, these operators rely on destructive scanning pipelines—slicing spines and feeding loose leaves through high-speed OCR rigs, reducing irreplaceable editions to industrial waste.
According to an investigation published on Anna's Blog by Anna's Archive volunteer "u", frontier labs routinely acquire massive stocks of out-of-print and secondhand volumes through proxy purchasing entities. A case in point highlighted in the report is Anthropic's confidential "Project Panama", which launched in early 2024 and surfaced during a $1.5 billion copyright settlement. Through this pipeline, Anthropic reportedly allocated tens of millions of dollars to acquire millions of physical books, slice them for Claude's training runs, and dispose of the remains.
The Economics of Data Monopolization
Destructive scanning is fundamentally an economic and strategic calculus rather than an engineering constraint. Guillotining book bindings is orders of magnitude faster and cheaper than non-destructive overhead capture. More critically, physical destruction permanently prevents competitors from acquiring the same source material while eliminating physical evidence that could complicate post-settlement copyright audits.
"After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers."
As volunteer "u" of Anna's Archive emphasized, this dynamic inverts the public mission of data democratization. Commercial labs pitch foundation models as tools that make human knowledge universally accessible, yet their acquisition mechanics systematically privatize historical print culture into proprietary model weights without releasing OCR corpora back to the public domain.
The Shadow Library Counter-Strategy
This contraction of physical supply arrives just as the industry runs out of unharvested, clean digital corpora. The un-digitized inventory surviving across independent bookstores, estate sales, and secondary markets represents the final high-density deposit of authentic human cognition.
In response, shadow libraries like Anna's Archive have initiated an emergency preservation effort. The network is currently coordinating distributed volunteers to locate, non-destructively digitize, and upload out-of-print books, niche journals, and rare archival holdings before corporate procurement sweeps the secondary market bare. By offering lifetime memberships and bounties, they aim to build an un-enclosed public corpus before frontier labs pulp what remains of the analog commons.