AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots

ai training databook destructioncopyrightdata sourcinglarge language models

In a troubling practice that some are calling 'AI book burning,' major technology companies are reportedly destroying millions of physical books to train their large language models (LLMs). This approach, which has gained renewed scrutiny in 2026, raises significant ethical and environmental concerns.


The Scale of Destruction


Industry insiders estimate that tens of millions of books have been pulped or incinerated over the past few years, primarily to obtain high-quality text for training generative AI. While digital scraping remains common, physical books offer unique advantages: they are often free of the formatting errors and paywalls that plague digital archives, and they provide access to out-of-print titles not available electronically.


Why Physical Books?


Despite the existence of vast digital libraries, companies like OpenAI, Google, and Anthropic have turned to physical book destruction for three key reasons:


  1. Copyright Evasion: Scanning and destroying physical copies allows companies to claim they did not 'copy' the material in the traditional digital sense, exploiting legal loopholes around first-sale doctrine and fair use. In 2026, several class-action lawsuits are contesting this practice.

    1. Data Quality: Physical books often contain clean, professionally edited prose that is ideal for training AI to produce coherent, human-like text.

      1. Access to Rare Works: Many valuable historical and scientific texts exist only in print. Rather than digitizing them with care, companies are opting for destructive scanning, which involves cutting off the spines and feeding pages through high-speed scanners.

      2. Environmental and Cultural Impact


        Environmentalists estimate that this practice has resulted in the destruction of over 150,000 trees since 2023, contributing to carbon emissions and waste. Culturally, librarians and scholars warn that irreplaceable works—including annotated editions, rare prints, and books with significant provenance—are being lost forever. 'We are witnessing a new form of informational vandalism,' says Dr. Ellen Marston, a digital ethics researcher at MIT. 'These companies are prioritizing speed over stewardship.'


        Legal and Ethical Responses


        In response to growing backlash, several European countries have proposed legislation to ban the destructive scanning of books for AI training. The European Union's proposed 'Cultural Heritage Protection Act' (CHPA), introduced in early 2026, would require companies to use non-destructive digitization methods and to compensate publishers and authors. Meanwhile, in the United States, the Copyright Office is investigating whether the practice violates the Moral Rights of authors.


        A Call for Sustainable AI


        As AI continues to demand more data, the question of how we source that data becomes ever more critical. Critics argue that destroying millions of books to feed chatbots is not only unethical but ultimately unsustainable. Alternatives being explored include synthetic data generation, data cooperatives that fairly compensate creators, and the use of public domain works. The choice, industry observers say, is between building AI on a foundation of destruction or one of respect for intellectual and cultural heritage.


        In the meantime, archivists are racing to digitize at-risk collections before the scanners arrive. Whether their efforts can keep pace with AI's insatiable appetite for text remains to be seen.

        via Decrypt AI

Related