Why AI Companies Are Cutting Up Books to Train AI Models

For years, the rapid advancement of large language models (LLMs) relied on a seemingly infinite, easily accessible resource: the public internet. Tech developers scraped billions of web pages, absorbing articles, forums, and wikis to feed their algorithms.

This massive data harvesting built the foundation of modern generative AI. But the digital landscape is changing.

As synthetic, machine-generated content floods websites and copyright lawsuits mount against indiscriminate web scraping, frontier technology firms are pivoting to a much older, more reliable medium. They are turning to physical books.

Recently, independent booksellers across Australia and Europe noticed a highly unusual market trend. Anonymous intermediaries began placing massive bulk orders for diverse catalogs, scooping up everything from cheap, mass-market paperbacks to obscure, out-of-print titles.

Unsealed court documents from copyright litigation involving major players like Anthropic have finally revealed the destination for these literary hauls. The books are not being collected for human readers or library archives.

Instead, they are being subjected to a severe process known as destructive scanning, physically dismantled to extract their linguistic value for artificial intelligence datasets.

The Race for Pristine, Pre-AI Training Data

The capability of any artificial intelligence model is inextricably tied to the quality of the data it consumes. When machine learning systems ingest low-quality, repetitive, or synthetic information, their generated outputs severely degrade.

Today, the internet is rapidly filling with AI-generated text. Training a new model on data produced by an older model creates a toxic feedback loop, a phenomenon researchers call “model collapse.”

To build smarter, more capable systems, developers desperately need vast quantities of clean, human-authored text.

Physical books represent an unparalleled goldmine for this specific requirement. Unlike a hastily typed forum post, a published book offers long-form, rigorously edited, and logically structured thought.

Crucially, titles published before the generative AI boom guarantee a complete absence of synthetic text. Data-sourcing vendors are actively marketing these older print catalogs to tech firms as untainted, premium datasets.

Acquiring commercial e-books directly involves navigating digital rights management (DRM) and negotiating expensive licensing deals with major publishers. Buying physical copies through third-party intermediaries effectively bypasses these initial corporate roadblocks.

Legal experts note that this exploits a highly debated gray area in intellectual property law.

Developers argue that digitizing a legally purchased physical book solely for back-end algorithmic training constitutes “fair dealing” or fair use, attempting to forge a controversial middle path between securing official publisher licenses and scraping illegal shadow libraries.

The Mechanics of High-Speed Destructive Scanning

Converting tens of thousands of bound paperbacks into machine-readable text requires industrial-scale efficiency. Archival organizations, such as the Internet Archive, rely on non-destructive digitization.

They carefully photograph open pages using specialized, vacuum-powered rigs designed to protect brittle bindings and preserve the historical artifact. AI developers, however, operate on entirely different incentives, prioritizing raw speed and data volume over preservation.

This demand for speed drives the practice of destructive scanning. The physical process is straightforward but absolute.

Technicians use heavy-duty guillotine cutters to shear off the book’s spine entirely, instantly transforming a bound volume into a stack of loose, disconnected sheets. These loose pages are then fed into automated, high-speed document feeders that can process hundreds of pages per minute.

By removing the binding entirely, the scanners capture perfectly flat images, eliminating the page curvature and shadow distortion that usually complicate traditional book copying.

These crisp, high-resolution images are immediately processed through advanced optical character recognition (OCR) software. The OCR translates the visual shapes of words into raw text tokens, allowing the neural network to ingest the data directly.

Once the text is successfully extracted and mapped into the AI’s training architecture, the physical remains of the book are simply thrown into a recycling bin or shredded.

This systematic destruction highlights a fundamental shift in how we value printed media. Books are increasingly being acquired not to transmit knowledge to a human mind, but to be sacrificed for their raw data, converting centuries of human expression into the statistical weights that power modern algorithms.

Source: Official The Indian Express, "Why AI Companies Are Buying Physical Books to Train Their Models"
Pradeepa Sakthivel
Pradeepa Sakthivel

Pradeepa is an AI Enthusiast and Technology Journalist covering AI News, AI Tools, Product Reviews, Industry Updates, and other developments in the rapidly evolving world of artificial intelligence.

Articles: 243