Amazon destroys rare books to train AI models
A hidden AirTag revealed that Amazon in Las Vegas is stripping the spines off rare books and scanning them for AI model development, sparking an ethical debate about the destruction of cultural heritage.
A hidden AirTag revealed that Amazon in Las Vegas is stripping the spines off rare books and scanning them for AI model development, sparking an ethical debate about the destruction of cultural heritage.
Amazon, a company that once started as an online bookstore, is now systematically purchasing rare books, destroying them, and using their content to train its own AI models. The discovery was confirmed when a bookseller hid an AirTag in one of the books from a bulk order, and the device was then tracked to Amazon's AI training facility in Las Vegas, as reported by 404 Media and Ars Technica.
At the warehouse marked VGT3, whose door is adorned with a logo of a Tyrannosaurus rex about to devour a book, a team of workers is tasked with stripping the spines off books and scanning the pages. Amazon declined to comment on 404 Media's findings, providing only a brief statement: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use."
The statement does not explicitly mention AI training, but Amazon is developing its own models considered "frontier" technology, requiring vast amounts of unique data to remain competitive with companies like Google, OpenAI, and Anthropic. Companies are currently guarding their training data closely to avoid losing their edge, and texts from rare books that are hard to find online provide exactly that advantage.
According to discussions on internet forums reviewed by 404 Media, VGT3 workers suggested that Amazon ran out of books to scan earlier this year. The shortage was so severe that workers feared the warehouse would close. At one point, inventory completely ran out, but the facility remains operational, and bulk orders of rare books continue to be delivered and systematically destroyed.
The 404 Media investigation further solidified the bookseller's theory that AI companies target books with ISBN numbers to ensure the maximum volume of unique works in their datasets. Workers are trained to scan barcodes or ISBN numbers before scanning the books themselves. As 404 Media notes, this "adds further credibility" to the theory that "AI companies are trying to methodically scan every printed book in the world by going through a list of ISBN numbers."
For booksellers, the money can be good, but the risk that their carefully curated collections end up in destructive scanning raises an ethical dilemma. AI companies skip the step of assessing book value, hunting for cheap, unique ISBN numbers to fill their lists. They often target older books of lower monetary value, such as those never translated from languages rarely used today or never popular enough to be widely distributed.
"These books may have historical value, intellectual value, sentimental value," said the bookseller who placed the AirTag to 404 Media. "AI companies don't care. They just want content as a pile of words strung together."
Scottish bookseller Derek Walker warned to the BBC that AI companies should distinguish between works that won't be missed and antique copies that are "the only known surviving copy of an edition." "It would be a much more significant problem if one such copy were purchased for destruction, having survived for so long," Walker said.
On Reddit, book lovers were divided over whether it's problematic for companies to destroy rare books for AI training. "You didn't intend to buy that old paper book," one user commented. Another argued that "an obscure book from the 1700s is now a museum piece and can reveal everyday things we didn't know."
The original Reddit poster emphasized that AI companies "will never share the content" of scanned books "because they don't want anyone else to train their AI" on the same works. One commenter added that AI models will only make available "distorted, censored, and paywalled fragments of ideas" from forgotten works. "Even if you love technology, you can admit that the concept of an AI literally eating books to become more powerful is pretty dystopian," wrote a person on the Reddit thread.
Michael Burry, a well-known critic, reportedly called the practice "evil incarnate." On the other hand, Amazon workers on forums described the job as a "nice" opportunity for those seeking monotonous work with flexible hours.
Companies like Amazon need unimaginably large amounts of text to train their large language models (LLMs). The internet is already largely exhausted, and rare books, especially those out of print or impossible to find online, offer a new source of coveted data.
These texts are particularly valuable because there is no chance that anything published before 2022 was written by an LLM. When models are trained on AI-generated text, they risk "model collapse," which occurs when the quality of a model's output degrades after the model ingests too much AI-generated text.