What happened

An AirTag hidden inside a rare book just confirmed what booksellers had suspected for over a year: Amazon is destroying rare books to train AI, tearing them apart to feed its large language models. According to an investigation by 404 Media, a bookseller planted a tracking device inside a book that was part of a bulk order, and the tag led straight to an Amazon facility in Las Vegas known as VGT3.

Inside that warehouse, a dedicated team rips books from their spines and scans every page — a process that permanently destroys the original copy. In a detail that has drawn particular outrage, the door of the VGT3 unit reportedly displays a logo of a Tyrannosaurus rex about to devour a book. Amazon has not denied the practice. When asked for comment, the company gave a boilerplate statement: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use." Notably, the statement never mentions AI training directly.

Workers inside the facility reportedly told online forums that Amazon ran dangerously low on source books earlier this year — at one point, the supply reportedly ran out completely, sparking fears the site would shut down. It didn't. Instead, bulk orders of rare, hard-to-find titles kept arriving, and the shredding continued.

Why it matters

This story matters far beyond the world of rare-book collectors. It's a window into how frontier AI labs are quietly solving one of the industry's biggest bottlenecks: running out of unique, high-quality training text. Common web data has largely already been scraped, so companies racing to build competitive models — Amazon among them — are turning to physical, out-of-print books that never made it online.

Workers at the facility said they were trained to scan ISBNs and barcodes before digitizing each book, which supports a theory booksellers had floated for months: AI firms are trying to systematically capture every printed ISBN in existence to maximize the uniqueness of their training corpus. That's a meaningfully different strategy than scraping public web text, and it signals just how far companies will go to gain a data edge in the AI arms race.

It also matters because rivals like Anthropic and xAI have publicly stated they do not train on rare or antique books — meaning Amazon's approach is neither industry-standard nor uncontested. Investor and commentator Michael Burry reportedly called the destructive scanning practice "evil incarnate," reflecting a growing backlash from authors, publishers, and cultural preservationists.

How to use it today

For entrepreneurs, marketers, and creators, this story is a useful reality check on data provenance. If you're building products on top of large language models, it's worth asking vendors direct questions: where does their training data come from, and does their usage policy respect copyright and licensing? The answers increasingly shape brand risk, not just ethics.

MyKreaTool AI chat — try ChatGPT, Claude and Gemini in one place. Free on MyKreaTool.Open the tool →

If your business publishes original content — books, guides, courses, or research — assume it may eventually be scraped or purchased in bulk for AI training unless you actively protect it through licensing terms, watermarking, or distribution controls. Reviewing your content's terms of use and metadata now is far cheaper than fighting a dispute later.

On the flip side, this also creates opportunity: businesses that can demonstrate transparent, ethically sourced AI tools have a real differentiator right now. If you want to experiment with AI without worrying about opaque data sourcing, tools like the free utilities at [mykreatool.com](https://mykreatool.com) let you test AI-assisted writing, image, and content workflows without committing to a black-box enterprise platform first.

Who benefits

Amazon is the clearest short-term beneficiary — access to rare, previously undigitized text could meaningfully strengthen its frontier models relative to competitors that have publicly ruled out this approach. In a market where OpenAI, Google, and Anthropic are all fighting for training-data advantages, unique text is a genuine competitive moat.

Secondhand booksellers experiencing bulk orders have also benefited financially, with 404 Media noting booksellers had a historic sales year partly driven by AI firms buying up inventory. Some individual workers at VGT3 described the scanning job as a decent, flexible gig — evidence that not everyone inside the system views it negatively, even as critics call it destructive.

Risks

The most obvious risk is cultural and historical loss: many of these books are rare or out of print, meaning destroying the physical copy could permanently erase editions, marginalia, or printings that can't be replaced. Unlike digitization methods that preserve the original (non-destructive scanning exists and is used by libraries), this approach trades preservation for speed and volume.

There's also legal and reputational risk for Amazon. Copyright law around AI training remains unsettled, and mass-purchasing books specifically to strip and scan them — rather than licensing content outright — could invite scrutiny from authors, publishers, and regulators. For any company building on Amazon's AI models, this creates downstream reputational exposure: customers and partners increasingly care how training data was sourced, and "we bought it commercially" may not satisfy critics or regulators for long.

Conclusion

The AirTag investigation turns a long-standing rumor into documented fact: Amazon is destroying rare books to train AI models, and it's doing so at scale, driven by a genuine shortage of unique training text. For business leaders and creators, the takeaway isn't just about ethics — it's about understanding where your AI tools' intelligence actually comes from, and making sure your own content and IP are protected in a landscape where physical books are now raw material for machine learning.