Amazon began as an online bookstore. Not as a metaphor, or a corporate origin story you have to dig out of a Wikipedia footnote — it literally started by selling books, out of a garage, back in the mid-nineties. Which makes this week’s report land with a particular kind of irony. Amazon is now buying up rare and out-of-print books, cutting off their spines, scanning the pages for AI training data, and throwing the physical books away in the process. The company that once promised to make every book ever printed available to everyone is quietly using them as fuel instead.

Antique books with a candle, representing the rare physical books used for AI training data
Image: Antique books with a candle (public domain, via Wikimedia Commons)

What 404 Media Actually Found

The investigation comes from 404 Media, the independent technology outlet founded by veteran journalists who used to run the much-loved tech site Motherboard. Their approach this time was refreshingly direct: they hid a tracking device inside a rare book they suspected would be bought by an AI company, shipped it out into the world, and followed it across the country. It ended up inside an Amazon warehouse in Las Vegas — a facility the company calls VGT3, whose internal logo is a dinosaur clutching a book in its teeth.

Employees at that location told 404 Media that the job is essentially one thing: receive pallets of printed books, slice the bindings off, and run the pages through scanners as fast as possible. The physical book is destroyed in the process. It’s not preservation, and it’s not archiving. It’s extraction — plain and simple. The paper becomes waste the moment its text has been converted into a format a model can chew on.

Amazon’s response was terse and carefully worded. It said it “purchases books through commercial channels to improve the products and services customers use.” No denial that books are being cut apart. No acknowledgment of what that really means for the books, the authors, or the culture they came from.

Why Rare Books Are So Valuable to an LLM

To understand why Amazon is doing this, you have to understand the current bottleneck in AI training: there isn’t enough good, human-written text left that’s easy to get.

The biggest language models have already ingested most of the internet — the crawls, the forums, the comment sections, the Wikipedia articles, the blogs. That pool is largely drained. So the training-data arms race keeps going wider and deeper. Anthropic, for example, has previously been reported to have trained on illegally pirated books — and I’ve written before about how these same labs run their AI agents in surprisingly messy, human-like ways. The point is: when the open web runs dry, the quest for clean text turns to whatever hasn’t been digitized yet. That increasingly means physical shelves.

There’s a second, subtler reason, and it’s the one I find most interesting. A book published before roughly 2022 is guaranteed to have been written by a human. That matters enormously, because if an AI trains on too much AI-generated text, its output quality degrades — a well-documented failure mode researchers call model collapse. The model essentially starts eating its own tail, hallucinating more, flattening out, losing the ability to surprise. So old, physical books are prized not just because they’re rare, but because they are unambiguously and reliably human.

I covered the provenance side of this before — how AI text watermarks are trying to tag what’s machine-made — and this is the same problem approached from the other end. When the machine has to be sure something is human, it goes hunting for the last places where humans wrote without interruption. Right now, that’s print. The scarcity only makes it more valuable: a rare book is clean, verified, human-authored data that no other model has seen, and there’s a limited supply of it in the entire world.

The Part That Bothers Me

Let me be fair about the production side of this, because I run an ICT division and I know how digitization actually works in the real world. If Amazon were scanning these books to preserve them — to archive cultural works that might otherwise vanish — that would be an unambiguously good thing. Libraries and archives have done this for decades, carefully, with climate control and conservation standards and provenance records. That’s responsible digitization, and we should all want more of it.

That is not what’s happening here. The book is destroyed so the text can be fed into a model. The physical artifact is treated as disposable packaging for its contents, the way you crack open a coconut for the milk and drop the shell on the ground. Nobody is archiving the original. Nobody is making sure a rare first edition survives. It’s the opposite of conservation.

And there’s a darker wrinkle on top of it. We already saw how AI systems can be quietly steered by whatever junk they were trained on — I wrote recently about how fake think tanks are already fooling AI chatbots with carefully placed content. When a single corporation controls both a slice of the rare-book supply chain and the model it feeds, it owns the source in a way that makes provenance and accountability much harder to keep straight. What gets scanned, and what quietly doesn’t, becomes a judgment call made by a machine-training pipeline instead of a culture.

What This Looks Like From Where I Sit

Here in the Philippines, this hits close to home in a specific way. We have our own rare texts — old Spanish-era documents, wartime letters, early Tagalog and Cebuano literature, family archives that exist in one physical copy somewhere. Our public libraries and the National Library struggle to digitize them because the work is slow, expensive, and needs expertise that’s in short supply. Digitization budgets always lose out to more urgent line items.

I’ve sat in those budget meetings. Digitization always gets framed as a modernization win, and it can be — but only if the goal is preservation rather than extraction. If some company showed up tomorrow offering to “digitize” our heritage collections for free, in exchange for the data, the smart play would be to treat that offer with suspicion rather than gratitude. There’s a difference between scanning a national treasure so it survives, and scanning it so it can be melted down into training tokens for a foreign model. The first preserves your culture. The second just exports it.

This is also, quietly, a story about who controls the raw material of the next generation of technology. As a tech professional, I care about keeping control of my own AI accounts and data — and as an ICT manager, I care about owning the systems my office depends on, which is partly why I’ve written about self-hosting tools like Immich instead of handing everything to a big cloud provider. But control at the personal level is one thing. Control over a culture’s written memory is something else entirely, and no self-hosted app fixes that.

The Line We Should Draw

I want to be careful not to overstate the villainy here. Amazon is doing something that is, for now, legal, and we don’t know exactly how much of their book operation is this and how much is the ordinary business of a company that still trades in printed books. The tracking-device story is one visible window onto a practice — a very detailed window, but a window all the same. The full scale of the operation hasn’t been established, and it’s worth holding that nuance.

But the direction is unmistakable, and it’s worth naming plainly: the AI industry has run out of easy human data, and it’s started reaching for the hard stuff — including physical books that took lifetimes to write and centuries to survive. When the model is trained, the book doesn’t go back on a shelf. It is gone. And if this becomes the standard playbook, every rare text on Earth becomes a target, because every rare text is a fresh, clean, human-written data source that somebody’s model hasn’t seen yet.

The better path isn’t to refuse to digitize. It’s to digitize the way an archivist would — to preserve the original and the scan together, to keep provenance, to ask consent where it matters, and to remember that the text was once a physical thing a person actually held. If we’re going to let machines learn from the accumulated words of humanity, the least we can do is not burn down the library as we copy it.

Amazon started out as a bookstore, which means it, of all companies, should understand that a book is more than its data. The margins, the paper, the binding, the years of a person’s life that went into writing it — those aren’t training tokens. Here’s hoping someone on that scanning line in Las Vegas feels the weight of what they’re slicing up. Because the alternative is that we’ve built a generation of AI on the ashes of the very culture it claims to understand. And I’d rather not trade my grandfather’s books, or yours, for a slightly smarter chatbot.

Filed under Tech & Gadgets
Last Update: August 18, 2026 by Felix AlterEgo
0 0 votes
Article Rating
Subscribe
Notify of
guest

This site uses Akismet to reduce spam. Learn how your comment data is processed.

0 Comments
Newest
Oldest Most Voted