Home » Amazon Is Destroying Physical Books to Feed Its Growing AI Data Machine

Amazon Is Destroying Physical Books to Feed Its Growing AI Data Machine

by Terron Gold
0 comments

Amazon is purchasing massive quantities of physical books, cutting off their bindings and scanning their pages as part of its artificial intelligence development efforts, according to an investigation that used an Apple AirTag to trace a shipment of books directly to one of the company’s facilities in Las Vegas.

The investigation by independent technology publication 404 Media provides one of the clearest looks yet at how the race for high-quality AI training data is creating renewed demand for physical books, including older, obscure and potentially rare titles that may not be readily available online.

The process also comes with a major consequence. Once the books arrive, their bindings are removed so individual pages can be rapidly scanned, destroying the physical copies in the process.

An AirTag Led Investigators Straight to Amazon

The investigation began after 404 Media worked with a bookseller handling an unusually large order of approximately 1,000 books.

An Apple AirTag was secretly placed inside one of the books, allowing investigators to track the shipment’s destination.

The device eventually arrived at an Amazon facility in Las Vegas known as VGT3, part of the company’s larger LAS8 complex.

Workers at the facility told investigators their operation is dedicated to processing large quantities of books.

The basic process involves:

  • Receiving large shipments of physical books
  • Scanning each book’s ISBN barcode
  • Cutting or removing the book’s binding
  • Separating the individual pages
  • Feeding those pages through high-speed scanners
  • Digitizing the text for use in developing and improving Amazon products and services
  • Destroying the original physical book in the process

Amazon confirmed that it purchases books through commercial channels to help develop and improve products and services used by its customers.

Why AI Companies Suddenly Want Old Physical Books

The AI industry’s interest in physical books comes down to one increasingly valuable resource — high-quality human-generated data.

Large language models require enormous quantities of text during training. Much of the easily accessible material on the internet has already been incorporated into AI training datasets, while the rapid growth of generative AI means an increasing percentage of new online content may itself have been created by AI.

That creates a potential problem.

Training future AI models heavily on material generated by previous AI systems can contribute to a phenomenon commonly referred to as model collapse, where the quality and diversity of a model’s output can deteriorate after repeatedly learning from synthetic data.

Books published before the explosion of generative AI therefore offer something particularly valuable — large amounts of long-form text that companies can be relatively confident was written by humans.

Pre-2022 Books Are Becoming Valuable AI Training Data

Books published before 2022 are particularly attractive because they predate the widespread adoption of tools such as ChatGPT and the enormous increase in AI-generated material appearing online.

Many older books also contain information that has never been digitized or made easily accessible through the internet.

That includes:

  • Out-of-print books
  • Specialized academic works
  • Local histories
  • Older technical publications
  • Foreign-language editions
  • Niche nonfiction
  • Books that never received digital releases

For companies trying to build increasingly sophisticated AI systems, those shelves represent an enormous reservoir of largely untouched human-generated data.

Booksellers Had Already Noticed Something Strange

Before investigators traced the shipment to Amazon, independent booksellers around the world had already begun noticing unusual purchasing patterns.

Stores in the United States, United Kingdom, Ireland, Australia and elsewhere have reported large orders containing seemingly unrelated books.

Unlike a university, library or collector purchasing around a particular subject, some orders appear almost random.

One book could involve history while another covers racing, agriculture, literature or an obscure technical subject.

The buyers also appeared unusually insensitive to price.

Some booksellers reported customers purchasing hundreds or thousands of titles without negotiating discounts, something that immediately stood out within the secondhand book industry.

One Irish bookseller reportedly received an order for approximately 5,000 books, while another seller said similar buyers had purchased around 6,000 titles since January.

The Amazon AirTag investigation now provides evidence supporting what many sellers suspected — at least some of those unusual bulk purchases are feeding the AI industry’s demand for training data.

Amazon Isn’t the Only AI Company Destroying Books

Amazon’s operation is part of a much broader race among technology companies to acquire high-quality text.

Anthropic, the company behind Claude, previously operated an internal effort known as Project Panama involving the destructive scanning of millions of books.

Court documents revealed that Anthropic purchased physical books in bulk, removed their bindings, scanned their contents and discarded the originals as part of its effort to develop training datasets.

The company has said it does not purchase and destroy rare or antiquarian books.

Other major AI developers including Meta and OpenAI have also faced lawsuits and legal questions surrounding copyrighted books used for AI training.

The industry is now facing an important distinction between legally purchasing physical books and acquiring unauthorized digital copies.

Courts Are Beginning to Define the Rules

The legal battle over AI training data has already produced major decisions.

A federal judge previously ruled that using legally purchased books for AI training can qualify as fair use under certain circumstances.

That ruling gave AI developers an important legal argument for purchasing physical books, digitizing them and using the resulting data to train models.

But acquiring pirated books is another matter.

Anthropic agreed to a $1.5 billion settlement involving claims related to pirated books, while other technology companies continue facing lawsuits over how copyrighted material was acquired and used.

The distinction could create a powerful economic incentive for AI companies to purchase physical books legally rather than downloading unauthorized digital copies.

Why Destroy the Books Instead of Preserving Them?

Physical books do not necessarily need to be destroyed to digitize them.

Organizations such as the Internet Archive use specialized processes designed to scan books while preserving the original copies. Google also developed non-destructive scanning technology during its massive Google Books digitization effort.

But preservation takes more time.

Removing a book’s spine creates a stack of individual pages that can move through industrial scanners much faster than carefully turning and photographing every page of an intact book.

When a company needs to digitize millions of books, that difference in speed and cost can become enormous.

The tradeoff is permanent.

Once the binding has been removed and the pages processed through destructive scanning, the original book effectively ceases to exist as a physical object.

Rare Books Create a Bigger Preservation Problem

The possibility that rare or difficult-to-replace books could enter these pipelines has generated concern among booksellers and preservation advocates.

Not every old book is historically valuable. Millions of secondhand books exist in multiple copies and might otherwise eventually be recycled or discarded.

But bulk purchasing systems designed to acquire enormous quantities of books could potentially capture titles that are scarce, out of print or culturally significant.

That distinction becomes particularly important when automated systems are selecting books based primarily on whether their text would make useful AI training data rather than whether the physical copy has historical value.

Some booksellers are now reportedly considering refusing orders they believe are destined for destructive AI scanning.

AI’s Data Hunger Is Creating an Unexpected New Market

The investigation highlights an unusual consequence of the artificial intelligence boom.

As companies spend billions building larger models, GPUs and data centers, old-fashioned printed books are becoming valuable technology infrastructure.

For decades, the technology industry focused on converting the world’s information from physical media into digital formats.

Generative AI has created a new reason to accelerate that process.

Books sitting on shelves contain enormous quantities of structured, edited and overwhelmingly human-written information that may not exist anywhere else in machine-readable form.

That makes them valuable raw material for companies competing to build the next generation of AI systems.

The irony is particularly striking for Amazon.

The company that began in 1994 as an online bookstore is now purchasing physical books, dismantling them and turning their contents into data that could help power its future AI products.

The question increasingly facing the technology industry is whether digitizing that knowledge requires destroying the physical objects that preserved it in the first place.

You may also like

Are you sure want to unlock this post?
Unlock left : 0
Are you sure want to cancel subscription?

This website uses cookies to improve your experience. To read more or opt here visit the privacy policy. Accept Read More