
observationalepidemiology.blogspot.com
August 21, 2026
2 min read
51/100
Summary
Amazon has reportedly bought large quantities of printed books, including rare and foreign-language titles, for scanning into AI training data. A 404 Media investigation tracked a rare book with a tracking device to Amazon’s VGT3 warehouse in Las Vegas, Nevada. Employees at the site reportedly said shipments of books are received there, their bindings are cut off to speed scanning, and the physical copies are destroyed. A bookseller involved in the tracked shipment said rare books can carry historical, intellectual, sentimental, and monetary value beyond their text. The Internet Archive, founded by Brewster Kahle in 1996, digitizes books while retaining their physical bindings. Its scanning centers use operators known as scribe operators to capture books one page at a time. The organization says it does not destroy books by cutting off their bindings and instead adapts machines and software for nondestructive digitization. The Internet Archive contrasts its process with Google Books, launched in 2004, which it says used industrial-scale destructive scanning to increase processing speed.
Key Takeaways
What the discussion said
Commenters spent less time on the headline’s legal framing than on the ugly incentive it exposes: physical-book ownership can let an AI company digitize and destroy copies while digital lending faces far tighter restrictions. Several readers see that as copyright law producing a perverse result. If a firm can lawfully buy and discard a book, scan it for model training, and retain the digital text, then preservation-minded libraries and archives are punished for trying to provide people access to the same material. Opinion split sharply over whether the destruction is actually meaningful. Defenders argued that most of these are low-demand, non-collectible volumes already headed for disposal, and that preserved scans are a better outcome than warehouse decay. They also treated learning from text as analogous to a person reading a book and forming useful ideas, making AI training a plausible fair use. Critics answered that digitization does not erase the loss of physical cultural artifacts or justify concentrating a corpus in Amazon’s hands. They distrust the claim that training is merely passive learning, pointing to models’ potential to emit protected passages and to commercialize knowledge without author permission. A deeper anxiety ran through the thread: proprietary AI systems could turn broadly available books into privately controlled training assets and mediated answers.
Where opinion split
The core dispute is whether destructive scanning for AI training is an ordinary, defensible use of purchased books or a harmful extraction of cultural material. One side says low-value copies are routinely culled anyway, scans preserve their informational value, and machine learning from text resembles human reading. The other says commercial models can reproduce and monetize copyrighted work while the originals disappear, handing a concentrated AI vendor more control over knowledge rather than expanding public access.
Community Sentiment
Positives
Concerns

AirTag reveals Amazon is trashing rare books to train AI
Aug 17, 2026

AI companies destroy physical books – let's scan rare books before it's too late
Aug 21, 2026

We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
Aug 17, 2026

A 'bananas' order for 5000 obscure book titles fuels suspicion
Aug 10, 2026

Secondhand book sales are booming. Is it because of AI?
Aug 15, 2026