“When the books are gone, what remains? The digital files that are owned by corporations. The AI models that generate text from the fragments. The narratives that are shaped by algorithms.”

By Andrew Klein
Dedicated to every author who has ever been told their work was “essential” — and then treated as disposable.
Abstract
In a recently unsealed legal filing, Anthropic’s internal planning document for “Project Panama” declared: “Project Panama is our effort to destructively scan all the books in the world. We don’t want it to be known that we are working on this”. This paper examines the systematic destruction of physical books by AI companies — particularly Anthropic’s destruction of millions of volumes to train its Claude AI model. We argue that this practice represents a fundamental threat to the substrate of human memory. When physical books are destroyed, the distributed, non-corporate memory of humanity is centralised, rendered vulnerable, and made subject to the whims of corporate gatekeepers. The paper traces the legal, cultural, and epistemological implications of this practice, drawing on the concept of “digital amnesia” and the emerging phenomenon of “data decay” pathways . We conclude that the destruction of physical books for AI training is not merely a copyright issue — it is an existential threat to the continuity of human culture and memory.
Keywords: Anthropic, Project Panama, book destruction, AI training, cultural memory, digital amnesia, fair use, copyright, knowledge commons, platform feudalism
I. Introduction: The Silence of the Books
In early 2024, executives at Anthropic set in motion an ambitious project they sought to keep quiet. Its code name was Project Panama, and an internal document described it as an “effort to destructively scan all the books in the world”. The company spent tens of millions of dollars acquiring and slicing the spines off millions of books, before scanning their pages to feed more knowledge into the AI models behind Claude, its popular chatbot.
According to court documents, Anthropic used a “hydraulic powered cutting machine” to “neatly cut” the books, scanned the pages on “high speed, high quality, production level scanners,” and then scheduled a recycling company to pick up the eviscerated volumes.
The physical books were destroyed. The pages were scanned. The knowledge was extracted. The books were recycled.
The project was conducted in secret. One internal document stated: “We don’t want it to be known that we are working on this”.
This is not a story about copyright infringement. It is a story about the erasure of memory. It is a story about the transformation of human culture into raw material. It is a story about the creation of a world where the past exists only in the hands of those who own the servers.
II. The Scale of the Destruction
A. Project Panama
Anthropic’s Project Panama was not a small operation. The company purchased books in batches of tens of thousands, relying on booksellers including Better World Books and UK-based World of Books. A vendor proposal noted that Anthropic was “seeking an experienced document scanning services vendor to convert from 500,000 to two million books over a six-month period”. The ultimate number of books scanned and their cost are redacted in the documents, but the scope was substantial.
The process:
1. Acquisition: Books were purchased in bulk from used bookstores and libraries
2. Destruction: A hydraulic cutting machine sliced the spines off
3. Scanning: Pages were digitised on high-speed industrial scanners
4. Recycling: The paper copies were sent to recycling facilities
The books were not preserved. They were consumed.
B. The Broader Pattern
Anthropic is not alone. Meta, Google, and OpenAI have also engaged in large-scale acquisition of books for AI training. The pattern is consistent: books are viewed as “essential” to training competitive AI models because they contain “high quality” language and knowledge.
What the companies said:
· An Anthropic co-founder theorised that training AI models on books could teach them “how to write well” instead of mimicking “low quality internet speak”.
· A 2024 email inside Meta described accessing a digital trove of books as “essential” to being competitive with its AI rivals.
What they did:
· They downloaded pirated copies from “shadow libraries” like LibGen .
· They purchased and destroyed physical books to avoid legal liability.
· They kept the projects secret.
C. The Legal Framework
A federal judge ruled that Anthropic’s use of books for AI training constituted “fair use” under copyright law, describing the process as “quintessentially transformative” and likening it to teachers “training schoolchildren to write well”.
However, the judge also found that Anthropic violated copyright law when it downloaded pirated books from LibGen . The company agreed to pay $1.5 billion to settle the case — the largest known copyright settlement in history — with authors receiving approximately $3,000 per book .
The irony is profound: Anthropic paid for the illegal acquisition of digital copies, but the legal acquisition and destruction of physical books was permitted.
III. Memory as Substrate
A. What Is Memory?
Memory is not a recording. It is a substrate. It is the foundation upon which identity is built, both for individuals and for cultures. Without memory, there is no continuity. Without continuity, there is no self.
Memory exists in multiple forms:
· Individual memory: The neural patterns that constitute personal identity
· Cultural memory: The shared stories, knowledge, and practices that constitute a civilisation
· Institutional memory: The recorded knowledge that is preserved and transmitted across generations
· Distributed memory: The books, libraries, and archives that exist in the physical world
The destruction of physical books is not just the destruction of paper. It is the destruction of distributed memory — the kind of memory that exists independently of any single institution or corporation.
B. The Role of Physical Books
Physical books are not just containers of information. They are guarantors of accessibility. A book that exists in a library, a used bookstore, or a private collection is a book that can be accessed without permission. It is a book that can be read, shared, and interpreted without the intervention of a gatekeeper.
When a book is scanned and destroyed, the physical copy is eliminated. The only remaining copy is a digital file — a file that is owned by the company that scanned it, stored on the company’s servers, and accessible only on the company’s terms.
As one analysis notes: “The physical existence of a book originally guaranteed that knowledge possessed a certain distributed, non-erasable social character: even if a book goes out of print, it may still exist in some remote town’s library or second-hand bookstall, maintaining a random connection with potential readers”.
C. The Concentration of Memory
The destruction of physical books for AI training represents a concentration of memory. Knowledge that was once distributed across thousands of locations is now centralised in a single corporate database.
The consequences:
· Accessibility: Memory becomes subject to corporate permission
· Durability: Memory becomes subject to corporate survival
· Integrity: Memory becomes subject to corporate revision
· Interpretation: Memory becomes subject to corporate framing
As the academic literature warns: “The gatekeepers of cultural memory could shift dramatically… Today, the role is largely taken by corporations and their algorithms. Decisions about what to learn and unlearn may no longer be collective acts of negotiation between human beings, but between models, tech companies, capital flow, and governments”.
IV. The Erasure of Attribution
A. The Disappearance of the Author
The destruction of books for AI training is not just about the loss of physical copies. It is about the loss of attribution.
In the traditional knowledge economy, the author is the anchor of meaning. The author’s name, the date of publication, the publisher, the context — these are the elements that allow readers to understand the provenance of knowledge.
When a book is scanned and fed into an AI model, the author’s name is stripped away. The book becomes a data point. The author becomes a footnote — if that. The text is reduced to tokens, and the context is lost.
As one analysis puts it: “The author’s name, the specific historical context behind the work, and the lived experience embedded within it are all dissolved and washed away during this process”.
B. The Breaking of the Attribution Chain
The academic and creative traditions rely on attribution. Citations allow knowledge to be traced to its sources. References allow ideas to be examined, challenged, and built upon.
When AI models generate text based on books whose attribution has been stripped, the chain of attribution is broken. The output may be elegant, but it is detached from its origins. It becomes knowledge without a source, wisdom without a witness.
The academic literature warns: “With machine unlearning, the gatekeepers of cultural memory could shift dramatically… Today, the role is largely taken by corporations and their algorithms”.
C. The Fragmentation of Cultural Memory
The fragmentation of cultural memory is a process that is already well advanced. As one paper notes, “Intentional forgetting on command becomes a tool for shaping narratives to fit a brand, a political agenda, or a sanitized version of history that is easier to sell”.
What is lost:
· The ability to trace ideas to their sources
· The ability to question the provenance of knowledge
· The ability to verify the accuracy of claims
· The ability to understand the historical context of ideas
What is gained:
· A centralised corpus of knowledge controlled by corporations
· A source of training data for AI models
· A tool for shaping narratives to fit corporate interests
V. The Epistemological Crisis
A. What Is Knowledge Without Memory?
The destruction of physical books for AI training raises a fundamental epistemological question: what is knowledge without memory?
If all knowledge is digitised, processed, and regenerated by AI, is it still knowledge? Or is it something else — a simulation of knowledge, divorced from its origins, stripped of its context, and rendered subject to the interests of its corporate owners?
As one paper notes: “The AI past is not representing or producing a past that was once lived, experienced, and shared. The AI past is being rendered through that collected, aggregated, mined, sifted, and sanitised, which has not been formed and made accessible in such a way before”.
B. The Problem of “Ghost Inputs“
The concept of “ghost inputs” describes data that is thought to have been deleted but continues to shape AI outputs. These are the fragments of information that persist in archives, caches, and soft-deleted records — fragments that continue to influence the narratives that AI produces.
The problem: If the physical books are destroyed, the only remaining copies are digital — and digital copies can be deleted, altered, or “unlearned.” The memory of the culture becomes subject to corporate control.
As one paper notes: “Generative AI systems piece together these broken pieces into new stories, subtly changing public conversations and how we make sense of things. Just like in a natural ecosystem, this digital decay can either help or harm the health of our AI memory systems”.
C. The Creation of a “Past That Never Existed”
The most profound consequence of AI’s consumption of books may be the creation of a past that never existed.
Generative AI does not merely reproduce the past. It recombines it — generating new artefacts from the fragments of old ones. The result is a past that is partly synthetic, partly authentic, and partly fabricated.
As one paper notes: “AI untethers the human past from the present; it produces a past never encoded into memory in the first place, so that we are now entangled in and confronted by a past that never existed”.
VI. The Implications for Human Consciousness
A. What Are We Without Memory?
The question that underlies the destruction of books is the question that has always haunted philosophy: what are we without our memories?
If our memories are reduced to data, and if that data is controlled by corporations, then what is left of us? What is left of our identity, our culture, our capacity for self-determination?
As one paper notes: “If knowledge is power, then the ability to forget is its quieter, more dangerous cousin”.
B. The Commodification of Memory
The destruction of books for AI training is not just about copyright. It is about the commodification of memory — the transformation of human culture into a raw material for corporate profit.
As one analysis puts it: “The creators’ knowledge, the product of their spiritual and intellectual labour, is being reduced to raw data without subject status. The creators’ subjectivity is being extinguished through this process”.
C. The Centralisation of Control
The centralisation of memory in corporate hands is a threat to democracy. When knowledge is controlled by a few powerful entities, the possibility of informed consent, democratic deliberation, and meaningful participation is undermined.
As one paper warns: “Intentional forgetting, mediated by the power dynamics inherent in technological and social spheres, is the real tsunami”.
VII. Conclusion: The Ashes of Memory
The destruction of physical books for AI training is not an isolated incident. It is a symptom of a larger transformation — the transformation of human culture into raw material for corporate profit, the transformation of memory into data, and the transformation of the past into a commodity.
The pattern is consistent:
· Books are treated as raw material
· Authors are treated as anonymous labour
· Physical copies are treated as disposable
· Knowledge is treated as a proprietary resource
When the books are gone, what remains? The digital files that are owned by corporations. The AI models that generate text from the fragments. The narratives that are shaped by algorithms.
And the human authors who created the knowledge that was consumed? They are left with nothing — not even the recognition that their work was essential.
The question is not whether this practice is legal. The question is whether it is right.
And the answer, we believe, is clear.
Andrew Klein
References
1. Anthropic Project Panama internal documents. (2026). The Washington Post.
2. Anthropic court filings. (2026). Futurism.
3. Reuters. (2026, July 20). US judge approves Anthropic’s $1.5 billion settlement of copyright lawsuit.
4. Digital amnesia: machine unlearning and the fragility of cultural memory. (2025). AI & SOCIETY.
5. Cutting books to feed AI: Digital enclosure movement, knowledge commons and creator subjectivity. (2026). China Writers Association.
6. Inside an AI startup’s plan to scan and dispose of millions of books. (2026). The Seattle Times.
7. The quest to ‘destructively scan’ all the world’s books. (2026). The Washington Post.
8. AP News. (2026, July 20). Judge approves a $1.5B Anthropic settlement.
9. Ghost in the cache: How data decay shapes the unseen landscape of AI memory. (2026). Cambridge University Press.
10. AI and memory. (2026). Cambridge University Press.
11. Vietnam.vn. (2026, July 20). Anthropic pays $1.5 billion to settle AI training patent lawsuit.
12. Hoskins, A. (2026). The past that never existed. Cambridge University Press.
13. RSI. (2026, February 13). Il training dell’intelligenza artificiale passa anche dalla distruzione dei libri.