Story
August 18, 2026
Amazon’s AI book hunt puts rare volumes on the chopping block
A tracker hidden in a shipment of rare books led to an Amazon facility in Las Vegas where workers reportedly cut bindings for rapid scanning. Amazon says its book purchases improve customer products, while booksellers fear irreplaceable works are being reduced to training data.
A hidden tracker has turned a long-running fear among rare-book dealers into a tangible trail: books bought in bulk may be ending up in an Amazon warehouse to be physically dismantled for data.
The alarm had been building as sellers noticed unusual, no-haggle orders for obscure titles—exactly the kinds of books whose scarcity can matter more than their resale price. Then 404 Media placed an AirTag inside a shipment of rare books. It arrived at Amazon’s VGT3 facility in Las Vegas, where employees said they receive large shipments, “cut the bindings off” to speed scanning, and destroy the printed copy in the process.1
That finding sharpened the central dispute. Amazon did not specifically address AI training, saying only that it “purchases books through commercial channels to help develop and improve the products and services our customers use.”2 But the methodical bulk buying, barcode checks and high-speed scanning have fuelled booksellers’ belief that unique texts are being collected as raw material for frontier models.
The stakes are not limited to famous first editions. Dealers say overlooked foreign-language works, out-of-print academic titles and locally published books can carry historical or intellectual value even when they are cheap. A book that appears disposable in a database may be the last surviving copy of an edition.
Critics argue Amazon and other AI companies have a less destructive option. The Internet Archive’s preservation-oriented process scans fragile volumes page by page, while its staff say automated systems have not worked well for “brittle books, rare volumes, and other special collections.”3 That approach is slower and costlier—the opposite of an AI industry racing for vast new reserves of text.
Supporters of the data hunt can argue that scanning makes neglected material more useful, and some sellers benefit from bulk sales. Yet critics say access through a model is no substitute for preserving the original, particularly when companies keep their training collections private. Rare books are attractive precisely because they contain text unavailable online and untouched by AI-generated content, a premium resource in the scramble to build better models.4