BlackTree Security · Infrastructure · Automation · AI

Amazon Bought the Books. Ownership Does Not Settle the Moral Question.

An AirTag hidden inside a shipment of books has exposed something more important than the destination of a parcel. It has exposed how little the public knows about the physical supply chain behind artificial intelligence.

In an investigation published by 404 Media, journalists worked with an independent bookseller who had received an anonymous order for roughly 1,000 older and obscure books. A tracking device placed in one of the volumes followed the shipment across several US states and ultimately to an Amazon warehouse in Las Vegas. Employees told the publication that an internal operation removes bindings, scans pages and disposes of the physical books after digitisation.

The evidence strongly supports the existence of a destructive book-scanning operation. The reported link to AI training also deserves serious attention, although Amazon has not publicly identified the models, products or datasets involved. The company has said that it purchases books through commercial channels to help develop and improve products and services.

That distinction matters. It keeps the factual claim precise. It does not make the moral question smaller.

The easy response is that Amazon bought the books. Once a company has legally purchased a physical object, it can generally cut it apart, scan it or throw it away. A recent US federal court decision in the Anthropic litigation also found that converting lawfully purchased print books into digital copies for training could qualify as fair use in the circumstances before that court. Pirated copies were treated differently.

But legality is only the floor. The more difficult question is whether ownership of a copy gives a technology company an unlimited moral claim over the knowledge it contains.

A book is property, but it is also part of a record

A printed book is at least two things at once.

It is a physical object that can be bought and sold. It is also an instance of a cultural work produced by an author, an editor, a publisher and, often, a community of ideas. Older and obscure books may carry an additional value because few copies remain in circulation.

Destroying one common paperback is unlikely to impoverish the world’s cultural memory. Destroying thousands of scarce or out-of-print works at industrial scale may be different. The moral significance does not reside in a single severed binding. It emerges from volume, selection and opacity.

Were the books checked against library holdings before destruction? Were particularly scarce editions identified and preserved? Did anyone assess whether annotations, inserts, binding details or other physical features had historical value? Could the books have been scanned non-destructively, donated after use or offered to archives?

If the answer to these questions is unknown, that is already a governance problem.

Digitisation can preserve content, but a private training dataset is not the same thing as a public archive. A library makes material findable, attributable and available under defined rules. A proprietary AI dataset can absorb the same material while concealing what was collected, how it was processed and who may benefit from it.

When the physical copy disappears and the digital copy remains private, society may not have gained preservation. It may simply have transferred knowledge from a public market into a closed computational asset.

Buying a copy does not buy the relationship

The language of ownership can obscure the number of relationships involved.

The seller owns a particular copy. The author retains an intellectual and moral relationship with the work. Readers depend on access. Libraries and archives protect continuity. Publishers manage rights and distribution. AI companies seek data that can improve commercial systems.

Those interests are not identical, and a purchase transaction resolves only some of them.

This is especially relevant when books are acquired not to be read, resold or archived, but to be transformed into machine-readable material. The purpose changes the ethical character of the transaction. A buyer may be legally entitled to destroy a book, but an industrial actor extracting its content for model development has responsibilities that an ordinary reader does not.

Scale creates duty.

The larger the acquisition programme, the stronger the case for provenance records, scarcity checks, preservation rules and public disclosure. A company capable of buying and processing books by the thousand is also capable of designing safeguards before the first binding is removed.

The invisible labour and material cost of AI

Public debate often treats AI training data as if it exists naturally, waiting in a cloud to be collected. It does not.

Datasets are assembled through decisions about what to buy, copy, license, scrape, exclude, label and destroy. People perform that work. Physical materials are consumed. Cultural priorities are embedded. Legal interpretations are converted into operational policy.

The reported Amazon facility makes this supply chain unusually visible. Pages enter a scanner. A physical object becomes data. The original may then become waste. What is normally described as model development begins to look more like extraction.

That does not mean all scanning is immoral. Libraries, researchers and accessibility projects have digitised books for decades, often with enormous public benefit. The Internet Archive, for example, has documented scanning methods that preserve a book’s binding while creating digital access.

The relevant questions are who controls the resulting copy, whether the original survives, whether rights holders are respected and whether the public receives any durable benefit.

A destructive process for a closed commercial dataset occupies a very different moral position from preservation work conducted for public access. Treating both as mere digitisation erases the distinction that matters most.

The pre-AI premium creates a new pressure

Older books have become attractive partly because they largely predate the current flood of synthetic text. They offer language written by people before generative systems began filling the web with material produced by other models.

That makes the cultural record economically valuable in a new way. Human-created work is becoming a scarce input for systems whose output may further dilute the information environment.

There is an uncomfortable circularity here. Authors and publishers created the high-quality material. Technology companies can convert it into private training assets. The resulting systems then compete in markets for writing, research, illustration and education, including with the people whose work helped make the systems useful.

Even where a particular use is lawful, reciprocity remains unresolved. What is owed to creators whose work supplies the quality signal? What is owed to booksellers who believe they are serving readers rather than industrial ingestion pipelines? What is owed to future researchers if scarce physical copies disappear into a process that leaves no public catalogue?

These are not nostalgic objections to technology. They are questions about how value is created and distributed.

What responsible book acquisition for AI could look like

The answer need not be a ban on scanning. It should be a higher standard for companies operating at industrial scale.

At minimum, a responsible programme would disclose its purpose, maintain a verifiable inventory and distinguish common books from scarce or culturally significant editions. It would use non-destructive scanning where practical, preserve originals where replacement is difficult and offer usable books to libraries, archives or the second-hand market rather than automatically discarding them.

It would also document the legal basis for copying, provide meaningful routes for rights holders and explain whether the resulting material is used for training, retrieval, evaluation or another purpose. Independent oversight should test whether the company’s practice matches its public claims.

None of these measures would prevent AI development. They would simply recognise that a dataset is not ethically neutral because the acquisition budget cleared procurement review.

The deeper question is stewardship

The AirTag story attracts attention because it turns an abstract debate into a physical journey. A book leaves a seller, crosses the country and enters a facility where its binding may be removed. The knowledge continues in another form, but the object and its history may not.

The most important question is therefore not whether Amazon can do this. It may have strong legal arguments that it can.

The question is what a company should do when private capability meets shared cultural inheritance.

Does purchasing access to knowledge create only rights, or also responsibilities? Should rare material be treated differently from replaceable stock? Is a private corpus a form of preservation if nobody outside the company can inspect or use it? When AI developers consume the human record, what must they return to the society that produced it?

Technology leaders often describe AI as infrastructure for the future. Infrastructure earns legitimacy not only through innovation, but through stewardship. If companies want the public to trust systems trained on the accumulated work of human culture, they should be able to explain how that culture was acquired, how it was protected and who benefits after the scanner stops.

Buying the book may settle the invoice. It does not settle the moral account.

Sources and further reading

Leave a Reply

Your email address will not be published. Required fields are marked *