A Las Vegas Warehouse, a Box Cutter, and a Rare First Edition
Somewhere in a Las Vegas, Nevada, Amazon warehouse, employees receive large shipments of printed books, slice off their bindings to speed up scanning, and then discard what remains. The physical book does not survive the process. According to a report published by 404 Media and written by Emanuel Maiberg, Amazon has been acquiring massive quantities of books – including rare ones – scanning them as AI training data, and destroying them in the process.
404 Media tracked this directly: journalists identified a rare book they suspected would be acquired by an AI company for training purposes and followed it across the country to that Amazon facility in Las Vegas.
The team responsible for this operation apparently works under a logo featuring a dinosaur holding a book in its teeth.

A Practice the Industry Has Been Running for at Least Two Years
Amazon is not alone in this. What the industry calls “destructive scanning” – the physical destruction of a book in order to scan its contents – has been documented across multiple major AI companies for at least two years. The current wave of public attention began this past January, when The Washington Post broke the story that Anthropic had been running a secret internal program, reportedly called “Project Panama,” with the goal of destructively scanning books at enormous scale. Details about Project Panama surfaced during a copyright lawsuit that authors filed against Anthropic earlier in 2025.
That case made Anthropic look like an outlier. It is not. Active lawsuits against Meta, Google, Microsoft, and OpenAI now point to similar behavior across the industry. Amazon itself had previously denied practicing destructive scanning, or had at minimum stayed conspicuously quiet on the subject – until the 404 Media investigation made denial considerably harder to sustain.
What unites these companies is not just method but motive. Books, especially older printed ones, contain text that has never been indexed online. The internet, already scraped thoroughly, offers diminishing returns for training large language models. Physical books represent a different category of data: dense, editorially filtered, linguistically varied, and largely absent from any existing digital corpus. One Anthropic co-founder, according to documents The Washington Post uncovered in the Project Panama reporting, theorized that training AI on books could teach models “how to write well” rather than reproduce what he called “low quality internet speak.”

Why the Pre-2022 Cutoff Matters
There is a specific technical reason AI companies are particularly interested in books printed before 2022. Any text generated by an AI model that gets fed back into a subsequent model’s training data degrades that model’s output quality through a recursive process called “model collapse.” Books published before AI-generated text became widespread are, by definition, clean of that contamination. That makes them unusually valuable as training inputs – not just for their content, but for what they are not.
As Maiberg noted in his 404 Media report, printed books also arrive pre-organized. Unlike raw web scrapes, which require extensive filtering and cleaning, a published book carries its own structural logic. Paragraphs, chapters, argument, narrative – the formatting itself is part of what makes the data useful.
The particular focus on rare books, and on works by authors who are deceased, adds another dimension. Dead authors cannot join a lawsuit. They cannot speak to journalists or post on social media. Their estates may have legal standing, but the practical friction of pursuing corporate litigation is substantially higher than it is for living writers, who have already demonstrated – through the cases against Anthropic, Meta, Google, Microsoft, and OpenAI – that they are willing to fight. For AI companies, the math on dead authors is straightforward.
What Public Pressure Has – and Hasn’t – Done
Book destruction lands differently than other forms of data acquisition, and not only among people who read. The cultural symbolism is difficult to escape: a company feeding books into a scanner and sending the remains to a landfill reads, to much of the American public, as something close to desecration. Some AI companies appear to have registered this, adjusting their public posture or their methods – at least visibly – in response to the attention these stories generate. Whether internal practices have shifted is harder to know.
Amazon’s case presents a specific problem for that kind of soft accountability. Unlike Anthropic, whose Project Panama emerged partly through court documents, Amazon’s destructive scanning operation came to light because a journalist physically tracked a book through the supply chain and found it at a warehouse where employees described their entire job as cutting bindings and scanning pages. That is not a leak. That is an observable, ongoing industrial process.

The 404 Media report was published this morning. Amazon has not yet responded publicly to its specific findings. Somewhere in Las Vegas, the shipments are presumably still arriving.






