Artificial intelligence developers are increasingly pivoting away from digital web-scraping to source training material from physical used-book shelves as the internet becomes saturated with machine-generated content. This shift toward "dead-tree" media comes as researchers warn that training future models on AI-generated text leads to a degradation of quality known as "model collapse." By acquiring and digitizing physical books printed before the generative AI boom, companies aim to secure a "pristine" record of human thought, language, and specialized knowledge.
The trend gained significant attention following reports that ISBNdb, a major online book database, had begun offering a specialized service to source up to one million physical books per order for AI developers. The company’s marketing materials, which have since been removed, promised to find older, rare, and out-of-print titles tailored specifically to the needs of large language model (LLM) training. This secretive operation highlighted a growing industry practice: the "destructive scanning" of physical libraries to feed the insatiable data requirements of modern silicon.
Why AI companies are buying and destroying old books for training data
The primary motivation for this physical acquisition strategy is the preservation of data integrity. For years, AI models like GPT-4 and Claude were trained on massive crawls of the public internet, including Wikipedia, Reddit, and news archives. However, since the public release of ChatGPT in late 2022, the web has been flooded with AI-generated articles, social media posts, and product descriptions.
When an AI model is trained on data produced by another AI, it begins to lose the nuances of human expression and reinforces its own errors. This phenomenon, dubbed "model collapse" by researchers in the journal Nature, threatens the future development of more advanced systems. Physical books printed before 2022 serve as a guaranteed repository of human-only writing, free from the "slop" of modern chatbot-generated filler.
Furthermore, books provide a level of structural coherence that is rarely found in web data. A single book represents a curated, peer-reviewed, and edited body of work with a logical progression of ideas. AI companies are buying and destroying old books for training data because these volumes contain dense, authoritative information that has often never been digitized or indexed by search engines.
The mechanics of destructive scanning and industrial digitizing
The process of converting a physical library into a digital dataset is often a one-way trip for the books involved. While organizations like the Internet Archive use non-destructive methods to preserve the integrity of bound volumes, commercial AI developers often opt for "destructive scanning" to maximize speed and accuracy. This involves shearing off the spines of books with industrial guillotines to create a stack of loose-leaf pages.
These pages are then fed through high-speed scanners capable of processing thousands of sheets per hour. Optical Character Recognition (OCR) software converts the images into machine-readable text, which is then cleaned and formatted for inclusion in a training corpus. Once the digital copy is verified, the original physical remains—now a pile of disconnected paper—are typically recycled or discarded.
This industrial-scale destruction has raised ethical and cultural concerns, drawing comparisons to historical losses of knowledge. However, from a technical standpoint, the loose-leaf method allows for perfectly flat scans, which reduces the errors and distortions often found in traditional book-scanning methods where the curvature of the page near the binding can blur text.
Anthropic and the legal precedent for physical book acquisition
While many companies operate in the shadows, court records have shed light on the scale of these operations. Anthropic, the developer of the Claude AI, was revealed in federal court filings to have purchased millions of physical books for its internal research library. The company reportedly hired a former Google Books executive to spearhead the effort, which involved spending millions of dollars to acquire used books in bulk from distributors and retailers.
In the case of Bartz v. Anthropic, U.S. District Judge William Alsup ruled in June 2025 that the company’s use of these books was "transformative" and fell under fair use protections. Crucially, the court found that because Anthropic purchased the physical copies and destroyed them after creating a single internal digital copy, they were not creating an unauthorized surplus of the work. This "one-in, one-out" logic has provided a legal roadmap for other developers to follow.
The ruling, however, did not absolve Anthropic of all copyright claims. The company was forced to reach a $1.5 billion settlement with authors and publishers over the use of more than 480,000 pirated books downloaded from "shadow libraries." This massive fine has only increased the incentive for AI companies to buy and destroy old books for training data legally, as a physical purchase provides a much stronger defense against copyright infringement lawsuits than digital piracy.
Impact on the used-book market and independent sellers
The sudden demand for obscure titles has caused a noticeable ripple in the used-book industry. Booksellers in both the United States and Europe have reported a surge in unusual bulk orders. Charlie D. Becker, a second-generation bookseller in Houston, noted receiving orders for dozens of obscure titles, including 30-year-old software manuals and local travel guides, which had previously seen zero market interest.
In the Netherlands, antiquarian dealers reported receiving lists of thousands of ISBNs ranging from academic folklore studies to technical engineering texts. While these orders provide a financial windfall for struggling shops, many sellers expressed unease about the ultimate fate of their inventory. There is a growing fear among bibliophiles that AI companies may be inadvertently destroying the last surviving copies of rare or niche titles that were never widely distributed.
The lack of transparency in these acquisitions is a primary point of contention. Because many AI developers require strict non-disclosure agreements (NDAs) from their suppliers, it is difficult for the public or historians to know which parts of the human record are being "ingested" and subsequently pulped.
Cultural reactions and the "Library of Alexandria" comparison
The image of industrial blades slicing through the spines of thousands of books has struck a visceral nerve with the public. Social media platforms have been flooded with comparisons to Ray Bradbury’s Fahrenheit 451 and the historical burning of the Library of Alexandria. Critics argue that while the information may be preserved digitally within a private company’s servers, the cultural artifact of the book is lost to the public forever.
In response to the controversy, some tech leaders have attempted to distance themselves from destructive practices. Elon Musk announced that his SpaceXAI team would be instructed to preserve rare books and scan them "the hard way"—using non-destructive methods that keep the bindings intact. This suggests a growing divide in the industry between those who view books as disposable data containers and those who see them as heritage items.
The debate also touches on the "dead internet theory," the idea that the web is becoming an echo chamber of bot-generated noise. As AI companies are buying and destroying old books for training data, they are essentially mining the past to sustain the future of digital intelligence. This has led some authors to argue that their work is being used to build the very tools that will eventually replace them, a sentiment echoed by several plaintiffs in the Anthropic litigation.
The future of human-authored data reserves
As the supply of pre-AI human writing becomes more valuable, the market for physical media is expected to remain tight. Some industry analysts suggest that we are entering a "post-human" data era, where the only reliable source of new, high-quality training material will be found in physical archives, private letters, and historical records that have escaped the reach of the internet.
The pivot toward physical books highlights a profound irony of the digital age: the most advanced technology in the world is currently dependent on the most traditional form of information storage. For AI developers, the dusty shelves of a used bookstore represent the last frontier of "clean" data.
Whether this trend continues will depend on future legal rulings and the evolving capabilities of AI to filter its own output. For now, the hunt for physical books remains a cornerstone of the AI arms race. The legacy of these millions of scanned and discarded volumes will be written in the code of the next generation of chatbots, even as the physical pages themselves disappear from the world’s libraries.











