Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
ai-ethicsdigital-preservationllmscorporate-responsibility

AI companies destroy physical books – let's scan rare books before it's too late

AI companies destroy physical books — let’s scan rare books before it’s too late

annas-archive.gl

August 21, 2026

3 min read

🔥🔥🔥🔥🔥

59/100

Summary

Anna’s Archive volunteer “u” called for worldwide volunteers to scan and upload physical books, periodicals, rare works, and archival materials to shadow libraries before copies are lost or become inaccessible. The August 5, 2026 guest post says small contributors may receive recognition and lifetime Anna’s Archive membership, while the organization can help cover scanning costs and offer other rewards for large-scale book uploads. The post alleges that several AI companies have bought large quantities of secondhand books through intermediaries, scanned them for training data, and destroyed the originals. It says Anthropic launched a confidential initiative called Project Panama in early 2024, spending tens of millions of dollars to purchase and scan millions of paper books for Claude training before destroying them; the post says the project emerged in a $1.5 billion copyright settlement. The volunteer argues that destruction can prevent competitors from obtaining the same books, reduce legal exposure, and cost less than lossless digitization. The post further claims that AI-generated material has accounted for more than half of newly published internet content since the beginning of 2025, and warns that digital preservation of human-created works is becoming more urgent. It calls for contributions through scanning, purchasing material for scanning, or donations.

Key Takeaways

  • Anna’s Archive volunteer “u” is seeking volunteers to scan and upload books, journals, newspapers, magazines, rare books, and other archive materials to shadow libraries.
  • The post alleges that AI companies are buying, scanning, and destroying physical books to obtain pre-2022 training data and limit others’ access to digital copies.
  • According to the post, Anthropic’s Project Panama began in early 2024 and involved purchasing and scanning millions of paper books for Claude training before destroying them.
  • Anna’s Archive says it offers recognition and lifetime membership for small scans and uploads and may help fund large-scale scanning projects.

What the discussion said

The thread spent less time mourning individual volumes than interrogating the legal and economic machinery behind AI training corpora. Many commenters argued that book destruction is a sensational symptom of copyright rules: companies can cheaply acquire a physical copy, scan it for model training, and then discard it while avoiding the far murkier rights attached to digital editions. They stressed that trainers need only one copy of any title, whereas publishers, libraries, and used-book channels routinely pulp far larger numbers of unwanted books. Several also doubted that genuinely unique artifacts are involved, suspecting the targets are scarce but still replaceable works. Even so, readers saw an ugly concentration of power in private firms converting hard-to-access human knowledge into proprietary training assets. The appeal of pre-2022, human-authored text was understood as both a source of cleaner training material and a way to distinguish machine-generated content. Some saw LLM access as an improvement over rare books sitting unread in collections, while others insisted that model-mediated access is not public preservation. The proposed remedy was not simply outrage at AI, but durable public digitization: copyright reform, archival exemptions, and eventual release of scanned materials through institutions such as the Internet Archive. A smaller group rejected the premise that destruction is legally required, calling it a cost-saving business choice rather than an unavoidable compliance measure.

Where opinion split

The sharpest fight was over whether AI-driven book disposal is a serious cultural loss or a misleading anti-AI symbol. Critics argued that private model builders are swallowing inaccessible knowledge while discarding physical copies, making preservation and public access more urgent. Defenders replied that one scanned copy is negligible beside routine industrial pulping, likely does not mean one-of-a-kind books are vanishing, and reflects restrictive copyright markets more than an AI-specific assault on literature.

Read original article

Community Sentiment

Mixed

Positives

  • Pre-2022 human-written books offer comparatively clean training material, and that provenance can help models or filters detect synthetic text rather than recursively learning from AI sludge.
  • Training on a single scarce title can expose its knowledge through AI systems to far more people than a copy locked in a library or private collection.
  • The controversy is pushing concrete preservation ideas, including nonprofit digitization rights and eventual public release of AI-scanned works after copyright expires.
  • Several commenters see the episode as evidence that copyright reform could unlock broader, lawful access to training data and cultural archives instead of leaving them behind private legal walls.

Concerns

  • Destroying physical copies after digitization feels like private AI firms converting cultural memory into proprietary model fuel while the public receives neither the scans nor archival guarantees.
  • Copyright restrictions on electronic editions may perversely steer model builders toward buying, scanning, and discarding print books instead of using durable licensed digital copies.
  • Claims that disposal is legally necessary drew skepticism; if shredding merely trims processing costs, AI companies are choosing convenience over preservation.
  • Even if most titles have surviving copies, scarce works can become harder to independently verify or access when their digitized form is captured inside closed training datasets.

Related Articles

It is a sign of the times that Amazon gets to call this fair use while huge corporations try to sue the Internet Archive out of business.

It is a sign of the times that Amazon gets to call this fair use

Aug 21, 2026

A ‘bananas’ order for 5,000 obscure titles from a bookshop in Galway fuels suspicion

A 'bananas' order for 5000 obscure book titles fuels suspicion

Aug 10, 2026

Secondhand book sales are booming. Is it because of AI?

Secondhand book sales are booming. Is it because of AI?

Aug 15, 2026

Blocking the Internet Archive Won’t Stop AI, But It Will Erase the Web’s Historical Record

Blocking Internet Archive Won't Stop AI, but Will Erase Web's Historical Record

Mar 21, 2026