Report alleges Amazon destroyed rare books to obtain AI training data
Futurism reports that Amazon used a process that involved destroying physical books, including rare volumes, in order to digitize their contents for AI training. The article frames the practice as a data-sourcing and ethics controversy around how training material was acquired, rather than a new model, product, or financing announcement.
Why it matters: Training-data provenance is a growing fault line in AI, especially when companies rely on copyrighted, scarce, or culturally significant material. If substantiated, the reported practice would sharpen concerns about consent, preservation, and the methods major companies use to assemble competitive datasets.
Sources
- Amazon Caught Destroying Rare Books to Train AI Futurism · August 17, 2026
- Hidden Airtag reveals Amazon is trashing rare books to train AI Ars Technica · August 17, 2026
Related stories
Independent events that offer a meaningful comparison, without implying that one caused the other.
-
Twitch adds creator opt-out for training Amazon generative AI models
Twitch added a creator opt-out for future Amazon generative AI training use, while reporting said creator content had already been available for Amazon AI development under existing terms, creating an independent consent-focused training-data governance comparison.
Story history
Direct context and developments in this event’s history.
Later developments
-
Advocates file FTC complaint over alleged destruction of books for AI training data
The FTC complaint turns that reported book-destruction controversy into a formal consumer-protection and preservation matter, escalating scrutiny of alleged AI training-data acquisition methods beyond media reporting and public criticism.