Book Massacre Fuels AI Gold Rush

Row of server racks in a data center
Photo: Gorodenkoff / Shutterstock

As Silicon Valley chases better artificial intelligence, quiet book‑destroying operations are shredding our shared history for data.

Story Snapshot

  • AI firms are buying millions of pre‑2022 books, scanning them, and then destroying the originals for training data.
  • Internal documents from Anthropic’s “Project Panama” describe a plan to “scan all the books globally in a destructive manner.”
  • A U.S. court ruling that this book‑to‑data pipeline is “fair use” opened the door for similar mass copy‑and‑destroy schemes.
  • Rare book dealers in Europe warn that irreplaceable antique volumes are disappearing into AI supply chains with no preservation.

Silicon Valley’s Secret Book‑Shredding Machine

Reporting from multiple outlets shows that artificial intelligence companies have moved beyond scraping the internet and are now targeting physical books on a massive scale. Since around 2024, these firms have quietly bought huge lots of used and antique books from library surplus channels and secondhand marketplaces, then run them through destructive scanners that slice off the spines and feed loose pages into high‑speed imaging systems. After scanning, the paper is pulped or recycled; the digital text is kept as training data for large language models.

Court documents unsealed in a major copyright case against Anthropic reveal the clearest example of this pipeline. The company hired a scanning contractor and bought stock from resellers such as Better World Books, aiming for order volumes from hundreds of thousands up to millions of books at a time. Internal planning language called the effort “Project Panama” and described it as an attempt “to scan all the books globally in a destructive manner,” signaling that disposal of the originals was not an accident but part of the plan.

Why AI Labs Say They Need to Destroy Books

Industry reports explain that AI firms now talk about “data scarcity” and “model collapse” as the reason they are raiding old book markets. Online text has been scraped heavily and is increasingly polluted with machine‑generated content, which can make new models worse when they train on it. Pre‑2022 printed books, especially older technical and historical works, are seen as a clean, human‑edited source of language and knowledge that has not yet been flooded by artificial text. That makes them attractive, even if accessing them means cutting up and discarding the physical copies.

Companies also have a legal motive for destruction. In the Bartz v. Anthropic case, a federal judge ruled that buying books, scanning every page, destroying the physical copies, and keeping only the digital files for artificial intelligence training qualified as “fair use” under United States copyright law. That ruling effectively told technology firms that the safest copyright path is to turn purchased books into data and then eliminate the originals to show they are not redistributing copies. Critics argue this encourages a “data first, heritage last” mindset that ignores the cultural and historical value of the physical works.

European Dealers Warn of Cultural Loss and Opaque Buyers

Rare book dealers in several European countries say they now see mysterious buyers, often acting through intermediaries, requesting large, subject‑agnostic orders of out‑of‑print and antique titles at three to five times the usual market price. These buyers show no interest in condition or specific content and only care that the books are printed before 2022, which fits the artificial intelligence industry’s hunt for “slop‑free” human text. Associations of antiquarian sellers fear that once these volumes are shipped to logistics centers and scanning warehouses, they are shredded, with no public record of what was lost.

One investigation describes a Canadian firm, Zoom Books, bulk‑buying tens of thousands of old books from Spain and other countries, then sending them to a United States facility where the texts are scanned and the originals recycled. Sources linked this operation to artificial intelligence model training, noting that the pattern mirrors Anthropic’s earlier Project Panama pipeline. Dealers warn that legal, religious, and local history works that once lived in specialty shops are now disappearing into black‑box data centers, turning unique physical artifacts into proprietary digital fuel for private algorithms.

What This Means for Readers, Heritage, and Policy Under Trump

For everyday readers and families, these operations raise simple questions: who decides that a book’s highest use is to be shredded for software, and who speaks for the communities that value these texts as part of their heritage? The practice does not directly violate the United States Constitution, but it touches conservative concerns about overmighty corporations, lack of transparency, and the quiet erasure of cultural memory in the name of efficiency. Once a rare local law compendium or theology volume is destroyed, no future child can hold or study that exact artifact again.

Under President Trump’s administration, Congress and regulators now face a clear policy choice. They can leave decisions about destructive scanning to private firms and judges focused only on copyright, or they can push for guardrails that protect physical collections and require preservation when books have historic value. Some proposals from heritage advocates include mandatory non‑destructive digitization for rare works, public registries of volumes destroyed for artificial intelligence, and incentives for partnerships with libraries instead of opaque bulk buyers. For conservatives worried about elites rewriting the record, this fight over books and bots is one more front in the larger battle to defend memory, ownership, and common‑sense limits on high‑tech power.

Sources:

feedpress.me, futurism.com, washingtonpost.com, vreme.com, youtube.com, ohiosap.org, roger-pearse.com, xh.umi6.com, gadgetreview.com, instagram.com, facebook.com