AI companies are destroying antique books
Allegations that rare and out-of-print books are being purchased in bulk, scanned for AI models, and then destroyed are sparking backlash.
12punto
Findings that artificial intelligence companies are purchasing rare, antique, and out-of-print books in bulk to train large language models—and then shredding and destroying the physical copies after scanning them—have sparked a new debate among authors, archivists, and book lovers.
According to documents that surfaced in a lawsuit filed against Anthropic, the company planned to acquire millions of physical books as part of a process called “Project Panama,” cut off their bindings, run the pages through high-speed scanners, and then destroy the originals. The documents also indicated a goal to scan between 500 thousand and 2 million books under a six-month supplier contract.
Another element fueling the controversy is the claim that the practice is not limited to bestsellers or easily accessible books. In an investigation published by 404 Media, the movement of packages was tracked using an Apple AirTag placed inside an order of approximately one thousand books purchased from an unnamed bookseller. It was reported that the package traveled from California to Milwaukee, and from there to a warehouse coded VGT3, which is stated to belong to Amazon.
According to the report, employees at the facility described their work as “book scanning.” These findings have strengthened claims that major technology companies are conducting industrial-scale book scanning processes to generate AI training data.
COPYRIGHT AND DATA QUALITY ARE AT THE CENTER OF THE DEBATE
It is stated that there are two main reasons why companies are turning to physical books: to reduce legal risks and to access higher-quality training data. The use of pirated online archives had previously left AI companies facing copyright lawsuits filed by authors, artists, photographers, and news organizations.
Legally purchasing a physical book provides the right to dispose of that copy in the U.S. According to critics, this creates a legal loophole that allows companies to build large digital datasets without having to pay content creators additionally. In some court assessments, the digitization of legally acquired printed books for internal model training purposes can be considered under fair use.
However, the issue is not limited to copyright. The destruction of rare and out-of-print books is also drawing backlash regarding the preservation of cultural heritage. Critics argue that some books may have been printed in limited numbers and that the elimination of physical copies could cause irreparable losses for researchers, libraries, and collections.
From the perspective of AI companies, books published before 2022 offer texts that were written by humans and have undergone editorial processes. As the rapid proliferation of AI-generated content on the internet increases the risk of new models being trained on AI-generated text, printed books are seen as a cleaner and more consistent source of data for companies.
There is no definitive data on the number of books destroyed. However, court documents and bulk orders reported by booksellers show that orders can range from hundreds of books to the scale of millions. Therefore, the debate brings to the agenda a broader question about how the race for AI development affects not only the digital world but also physical cultural assets.