A coalition of more than a dozen civil society organizations has formally petitioned the Federal Trade Commission to investigate a controversial practice where artificial intelligence companies acquire physical books, digitize them for training purposes, and subsequently destroy the hard copies. The groups argue that this "hoard-and-destroy" methodology constitutes a violation of federal antitrust laws by systematically removing essential training materials from the reach of potential competitors.
The formal request, led by organizations including the Demand Progress Education Fund, the Consumer Federation of America, and the Institute for Local Self-Reliance, signals a new front in the regulatory battle over the future of generative AI. Advocates contend that by purchasing and destroying physical volumes, dominant AI firms are creating an "insurmountable systemic moat" that prevents smaller startups and researchers from accessing the high-quality data necessary to build competitive models.
The practice has drawn sharp criticism from both tech industry watchdogs and cultural preservationists, who liken the systematic destruction of physical media to modern-day book burning. While much of the legal debate surrounding AI has focused on copyright infringement and digital scraping, this latest development centers on the physical control of information as a means of market dominance.
The Rise of the "Hoard-and-Destroy" Strategy
As the race to develop more sophisticated "agentic" AI intensifies, the demand for high-quality, human-generated text has reached a fever pitch. Reports indicate that several prominent AI developers have turned to the physical book market to source data that has not yet been digitized or is otherwise difficult to acquire through standard web crawling.
Once these books are acquired, they are processed through high-speed industrial scanners. After the text is converted into a digital format suitable for machine learning, the physical artifacts are reportedly destroyed. This ensures that the digital copy remains a proprietary asset of the company, while the physical source material is permanently removed from the secondary market, libraries, and private collections.
Industry analysts suggest that this strategy is a response to the "data wall"—a point at which AI companies have exhausted the high-quality text available on the open internet. By targeting physical books, these companies are tapping into a massive, offline repository of human knowledge that remains largely untapped by current large language models.
Why Pre-2022 Literature is a Strategic Asset
The push for book destruction is particularly focused on volumes published before 2022. These works are considered "clean" data by AI researchers because they are guaranteed to have been authored entirely by humans without the assistance of generative AI tools.
Data scientists warn of a phenomenon known as "model collapse," where AI models trained on AI-generated content begin to produce nonsensical or degraded output. To avoid this recursive loop of degradation, companies are desperate for "virgin" human text. Books published in the mid-to-late 20th century are especially prized for their rigorous editorial standards, which provide a level of linguistic precision and factual density that is often lacking in modern web content.
By destroying these books after scanning, AI companies are effectively "salting the earth" for their rivals. If a dominant firm buys up the remaining physical copies of rare technical manuals, historical texts, or out-of-print literature and destroys them, a new competitor cannot simply go to a used bookstore or a library to find the same data.
Antitrust Concerns and the "Systemic Moat"
The letter sent to the FTC argues that this behavior is a textbook example of anticompetitive conduct. Under Section 5 of the FTC Act, the commission has the authority to investigate and prosecute "unfair methods of competition." The civil society groups argue that the destruction of books serves no legitimate business purpose other than to raise the barrier to entry for others.
"Unlike a standard data acquisition strategy, this hoard-and-destroy practice could serve as yet another structural mechanism to raise rival companies’ costs," the coalition stated in their letter. They argue that this prevents "fledgling competitors" from accessing the key source material essential to competing in the rapidly evolving AI marketplace.
The FTC, under the leadership of Chair Lina Khan, has shown an increasing willingness to challenge Big Tech firms on unconventional antitrust grounds. The commission has previously expressed concern over how dominant platforms use their vast resources to lock in advantages, and the destruction of physical resources to create a data monopoly could fall squarely within the agency’s current enforcement priorities.
The Shift Toward Knowledge as a Utility
The controversy over book destruction stands in stark contrast to the public rhetoric of AI leaders. OpenAI CEO Sam Altman has frequently compared AI-generated knowledge to a public utility, suggesting that in the future, intelligence will be sold to consumers "like electricity and water."
Critics point out the irony in this vision: while the industry promises to democratize information, its back-end practices involve the literal destruction of the world’s oldest and most reliable repository of human knowledge. Unlike electricity or water, which are renewable or cyclical, a destroyed out-of-print book is a permanent loss to the public record.
The move to turn knowledge into a proprietary utility also threatens the traditional role of libraries and the "first sale doctrine," which allows individuals and institutions to lend or resell books they have purchased. If the primary buyers of physical books become AI firms intent on destruction, the price of physical media could skyrocket, making it inaccessible to the public and educational institutions.
Cultural and Ethical Implications
Beyond the legal and economic arguments, the practice has sparked an ethical outcry. The systematic destruction of books carries heavy historical baggage, often associated with censorship and the erasure of cultural heritage. While the AI companies are not destroying books to suppress ideas, but rather to monopolize them, the result is seen by many as a strike against the "information commons."
The Consumer Federation of America and other signatories emphasize that the books being targeted are often those that are not easily found in digital formats. This includes local histories, specialized scientific treatises, and cultural works from marginalized communities. If these are destroyed after being scanned into a private corporate database, the public’s ability to verify the information or use it for non-commercial research is severely limited.
Furthermore, the environmental impact of industrial-scale book destruction adds another layer of scrutiny. The energy-intensive process of training AI models is already under fire; the added waste of destroying millions of physical pages contributes to a growing perception of the industry as being indifferent to physical-world consequences.
Potential Regulatory Outcomes
If the FTC chooses to act on the petition, it could result in a series of mandates or lawsuits aimed at curbing the destruction of training materials. Potential remedies could include a requirement for AI companies to donate books to libraries or archives after scanning, or a "right to compete" provision that would force companies to make their training datasets available to others if the original sources were destroyed.
The investigation would likely involve subpoenas to major AI developers to determine the scale of their physical acquisition programs. Investigators would seek to understand the internal decision-making processes that led to the destruction of books and whether there was a documented intent to disadvantage competitors.
Legal experts suggest that this case could redefine how "assets" are viewed in the digital age. If a physical object’s primary value is the data it contains, the destruction of that object to prevent others from accessing that data could be seen as a new form of predatory behavior.
The Road Ahead for AI Regulation
The push for the FTC to sue AI companies over book destruction is part of a broader global movement to rein in the power of artificial intelligence developers. From the European Union’s AI Act to various state-level initiatives in the U.S., lawmakers are struggling to keep pace with the speed of technological change.
The civil society groups are calling for immediate transparency. They are urging the FTC to require companies to disclose which titles they have acquired and destroyed, and to provide an accounting of how much of their training data comes from sources that are no longer available to the public.
As the FTC reviews the petition, the debate over the preservation of physical knowledge versus the advancement of digital intelligence continues to intensify. For many, the outcome of this struggle will determine whether the future of AI is built on a foundation of shared human heritage or on the ashes of the very books that made that intelligence possible.
The coalition concludes that a failure to act would set a dangerous precedent, allowing the wealthiest corporations to buy up and eliminate the physical evidence of human culture in the pursuit of a digital monopoly. The case now rests with federal regulators, who must decide if the protection of competition requires the protection of the printed word.












