A 3,001-Book Order That Wasn’t Spam

It was a Monday morning in July 2026 when an unusual email reached Pieter de Vries at his antiquarian bookshop in Haarlem, Netherlands. The sender sought what they called a “fairly large order” and attached a spreadsheet containing 3,001 ISBNs – mostly titles published between 2020 and 2021 by academic heavyweights like Elsevier, Wiley, Routledge, and Oxford University Press. The instructions asked de Vries to price the books, estimate shipping to China, and reply. He barely glanced at the attachment. “I regarded the request as spam or phishing,” de Vries told Fortune. His business, after all, deals in rare and expensive volumes – one book at a time, never in bulk.

Weeks later, a Dutch journalist investigating the procurement request showed him the same spreadsheet and explained its likely destination: artificial intelligence. It was a revealing moment. The email never mentioned AI, but it surfaced just as AI companies are turning to the physical world for high-quality text. Court records in a settled US case detail how Anthropic, through its “Project Panama,” bought millions of books, removed their bindings, scanned the pages, and discarded the originals to build a digital library for training its Claude models. The process, called “destructive scanning,” slices spines so that pages can be fed through high-speed scanners. An Anthropic spokesperson said the company does not acquire or destroy rare books and purchases through regular commercial markets, but the fact pattern – bulk orders of recent academic titles shipped abroad – mirrors what de Vries and several other Dutch antiquarian booksellers received. Similar reports have emerged from Switzerland, where secondhand sellers were approached with bizarrely large requests for specialized works.

While de Vries’s shop was never in danger of filling the order, the episode exposes a new layer of AI’s data pipeline: a physical supply chain that turns books into training data before the paper even has a chance to yellow. According to the investigative reports, a metadata company called ISBNdb considered launching a service to help AI labs source “1,000 to 1 million books” for LLM training, though it says the idea never went beyond an exploratory webpage. Regardless, the market signal is unmistakable: the web is no longer enough, and the world’s printed academic knowledge is becoming a commodity for large language models.

The Quiet Logistics Behind AI’s Text Appetite

Why AI Labs Are Buying Up Books Instead of Just Scraping the Web

The web is vast but noisy. Academic books, especially from top-tier publishers, offer dense, carefully edited, domain-specific knowledge that is gold for training models on medicine, law, engineering, and business. A single 400-page volume can contain more coherent, high-quality prose than thousands of forum posts. As AI companies push into professional and scientific applications, the value of this curated material skyrockets. Destructive scanning is also faster and cheaper than manually photographing or waiting for all titles to appear in digital form, many of which are locked behind expensive library subscriptions. The request that reached de Vries – targeting titles published only a few years ago – suggests the demand is for current, field-specific knowledge that isn’t easily acquired through existing digital licenses.

Advertisement

The Legal Gray Zone: Fair Use vs. Destroyed Copies

The Anthropic settlement established that using lawfully purchased physical books to train AI models is fair use under US copyright law. However, the ruling focused on the scanning and training use, not on the destruction of the physical copies. French, German, and other European copyright regimes do not have an exact equivalent of American fair use, and the systematic destruction of printed works could draw moral-rights claims from authors or trigger new regulatory scrutiny, especially when the books are out of print or culturally significant. The fact that the Haarlem inquiry came with no disclosure of the intended use adds another layer of risk: if bulk buyers disguise their purpose, booksellers may inadvertently become a conduit for practices that, while currently legal, could quickly become controversial.

Who Gains and Who Gets Sidelined

The winners are the AI labs that obtain large volumes of structured, authoritative text at a fraction of the cost of negotiating individual licenses. Large-scale secondhand book dealers may see a new, if peculiar, revenue stream – though the reputational cost of being seen to supply a book-destroying pipeline is real. Antiquarian sellers like de Vries are unaffected because their inventory is too scarce and expensive. For major publishers – Emerald, Elsevier, Wiley, and Oxford University Press are all named in the 3,001-entry spreadsheet – there is a double-edged sword: they sell more physical copies, but those copies then become the raw material for AI models that could compete with their own digital products. Authors and copyright holders, meanwhile, may find their work ingested into training sets without any additional payment, raising fresh questions about creator compensation in the AI era.

What the Shadow Book Trade Means for Publishers, Booksellers, and AI Labs

  • For publishers: Monitor secondhand marketplaces for bulk orders of your recent academic titles. Consider embedding licensing language in future agreements that addresses whether a physical sale grants the right to digitize and use the content for AI training.
  • For booksellers handling large inventory: When you receive an unsolicited order for thousands of current titles from an unknown buyer, verify the purchaser’s identity and intended use. A simple inquiry can help you avoid facilitating a data pipeline that could later attract legal or public backlash.
  • For AI developers: The days of opaque data sourcing are numbered. Even if a fair use defense holds in key jurisdictions, the optics of destroying books to feed a model can provoke regulatory responses and erode public trust. Proactive transparency about data provenance and a clear policy on physical-book digitization are becoming prerequisites, not optional extras.

Risk & Opportunity Assessment

Commercial RiskMediumIf public backlash or new legislation restricts bulk-book scanning, the cost and availability of high-quality training data could rise, disrupting data supply chains for AI labs.
Competitive RiskLowDestructive scanning is not yet a source of unique competitive advantage; multiple AI firms are exploring the same method, and the practice does not confer a proprietary data moat unless combined with exclusive sourcing.
Regulatory RiskMediumThe US settlement provides a fair use shield for purchased books, but European courts and legislators may impose stricter rules on the destruction of physical cultural goods and the opacity of bulk orders, especially if authors’ moral rights are invoked.
Reputation RiskHighThe imagery of companies slicing spines and discarding thousands of academic books is likely to provoke strong public and author opposition, potentially damaging brand trust for the AI firms involved and any intermediaries that facilitate the trade.
Technology DisruptionLowThe core technology—high-speed sheet-feed scanning—is well established. The disruption is in the data procurement model, not a breakthrough in digitization hardware.
Commercial OpportunityHighAcquiring millions of recent, professionally edited texts at low cost offers AI companies a direct path to higher-quality training data in specialized domains, which can drive model accuracy in medicine, law, and engineering—markets worth billions.