Yes, the clearest documented case is Anthropic, which bought millions of physical books in 2024, cut off their bindings, scanned them into digital files, and discarded the paper originals. The strongest primary-source evidence is a June 23, 2025 order in Bartz v. Anthropic, later expanded by a July 17, 2025 class-certification order and a January 27, 2026 Washington Post investigation.
What is documented is narrower than “the whole AI industry,” but it is not vague. Court records say Anthropic shifted in 2024 from building a library with pirated ebooks to buying print copies “in the millions,” then sending them through a destructive scanning workflow because physical books were a legally safer way to feed its growing need for training text amid AI’s data bottleneck and broader AI training data limits.
Anthropic is an AI lab best known for the Claude chatbot. In the court record, the company’s internal project was called “Project Panama”, and the point was simple: buy books, turn them into machine-readable text, and use that text to train models.
Anthropic’s Project Panama bought and destructively scanned millions of books in 2024
Judge William Alsup’s June 23, 2025 order states that Anthropic “purchased millions of print books”, then cut the bindings off, scanned every page, extracted the text into digitized, searchable files, and threw away the paper copies. That is the core answer. It was not an industry rumor. It was described in a federal court order.
The Washington Post’s January 27, 2026 investigation added the missing operational detail. It reported that Anthropic’s “Project Panama” spent “many millions of dollars” buying used books and working with a scanning vendor that proposed a workflow for 500,000 to 2,000,000 books. Some exact totals and costs were redacted in court filings, but the scale is still clear: this was an industrial pipeline, not a one-off digitization job.
The July 17, 2025 class-certification order fills in why Anthropic did it. The order says the company had already assembled a giant internal library of books, including pirated copies, and then in 2024 moved to buying print books because it wanted more material and a different legal footing. Physical books solved two problems at once: they were plentiful in the used market, and ownership of a lawful copy gave Anthropic a cleaner argument for making training files from them.
That pipeline appears to have worked roughly like this:
- books or lots of books were bought through used-book channels
- their bindings were cut off so pages could move through high-speed scanners
- the scans were converted into text and stored as searchable digital files
- the paper books were discarded after scanning
- the resulting corpus was used in model training
The ugly part is physical, not abstract. Destructive scanning is a real archival technique: if the goal is cheap, fast throughput, cutting the spine off a book is the equivalent of taking a document feeder to it. For a mass-market paperback, that is one thing. For an obscure out-of-print edition that may not be widely held, it is another.
What the public record does not establish is a verified count of how many destroyed copies were genuinely rare in the strict bibliographic sense. The documented record is strongest on scale and method, weaker on the exact rarity profile of the books being destroyed.
The court drew a line between lawful book purchases and pirated shadow libraries in 2025
The June 23, 2025 ruling did not say all of Anthropic’s book practices were fine. It drew a bright line.
Judge Alsup held that Anthropic’s conversion of lawfully purchased print books into digital files for internal search and AI training was fair use. The order treated the scanning as an intermediate use tied to analysis and training, not as a market substitute for the books themselves. On purchased print copies, Anthropic won that point.
But the same litigation record says Anthropic had also downloaded books from pirated “shadow libraries”. On that issue, the court was much colder. The June 23 order and July 17 order make clear that buying a physical copy and scanning it is legally different, in that court’s view, from simply acquiring pirated digital books in bulk.
“Anthropic had no entitlement to use pirated copies for its central library.”, Judge William Alsup, Bartz v. Anthropic, June 23, 2025
That distinction matters because it turns a broad culture-war claim, “AI companies are stealing books”, into a more specific factual map. One pipeline ran through lawful used-book purchases and destructive scanning. Another ran through piracy. The court treated them differently.
It is also not the last word for all AI training disputes. The June 23, 2025 decision is a U.S. district court ruling, not a final nationwide rule, and other copyright cases can come out differently on different facts.
The National Library of the Netherlands highlighted that exact point in a feasibility study for a European Books Data Commons. It noted that Anthropic’s litigation records offer an unusually direct window into how AI firms now source book data, unusually direct because companies almost never describe this pipeline voluntarily.
European booksellers’ 2026 reports widened the controversy from copyright to cultural preservation
By 2026, the controversy had moved beyond whether scanning purchased books could qualify as fair use. Booksellers and libraries started asking a different question: what kinds of books are being pulled into these bulk-buying pipelines?
A June 25, 2026 report from NL Times, citing European booksellers, said dealers were receiving unusual bulk requests for obscure and often rare books that they suspected were headed for AI training. A May 2026 trade report in Publishers Marketplace said Canadian company Zoom Books was buying out-of-print books for AI scraping. A July 21, 2026 El País report linked secondhand-book purchases, intermediaries including Zoom Books and PrepFort, and AI firms’ search for cleaner pre-LLM-era text.
A Tagesschau report from July 2026 pushed the same pattern into the German antiquarian trade, describing unusual orders routed through Zoom Books and shipped via Canada. None of that is as direct as the Anthropic court record. It is investigative reporting and bookseller testimony, not a lab filing an exhibit saying “we destroyed these exact editions.” But it widens the picture from one company’s legal defense to a live procurement market.
The preservation concern is practical. If a book is merely old, obscure, or out of print, destructive scanning is unpleasant but not necessarily a cultural loss if many copies survive. If a book is scarce, locally held, or special-collection material, destructive scanning can erase part of the remaining physical record. Public reporting has not verified how many destroyed copies fell into that second category. The fear is credible; the count is not yet nailed down.
Libraries are reacting as if the risk is real. A Statement of Shared Practice on AI Training Requests and Unique Cultural Collections, published by library-sector participants in 2026, shows institutions formalizing responses to AI training requests involving rare and special collections. That is a sign the issue has moved from internet rumor to collection-management policy.
The simplest way to understand the whole episode is that AI labs have run into data quality problems. Web text is abundant, but it is messy, increasingly synthetic, and often already contaminated by previous generations of models. Older books are attractive because they are edited, long-form, and mostly pre-LLM. That makes secondhand bookstores look less like quaint retail and more like unindexed data warehouses.
The next factual milestone is likely to come from either further unsealed litigation records in Bartz v. Anthropic or direct disclosures tying specific 2026 European sourcing networks to named AI labs.
Key Takeaways
- Anthropic is the clearest documented case of an AI company buying and destructively scanning physical books for training.
- Court records say Anthropic purchased and scanned books “in the millions” in 2024, though some exact totals and costs were redacted.
- A June 23, 2025, U.S. district court ruling treated scanning lawfully purchased print books differently from acquiring pirated books from shadow libraries.
- Public reporting has not established a verified count of how many destroyed books were truly rare rather than simply obscure or out of print.
- European bookseller and library responses in 2026 show the controversy has widened from copyright into preservation and collection-management concerns.
Frequently Asked Questions
Are AI companies really cutting apart physical books?
Yes, at least one major AI company is concretely documented doing it: Anthropic. The court record says it bought print books, cut off the bindings, scanned the pages, extracted the text, and discarded the originals. Broader claims about multiple companies are more weakly documented than the Anthropic case.
How many books did Anthropic scan?
The public court record says “millions” of print books. The Washington Post reported a vendor workflow for 500,000 to 2,000,000 books and said Anthropic spent many millions of dollars, but some exact counts remain redacted.
Did Anthropic destroy rare books?
That is not proven at scale in the public record. Reporting from European booksellers suggests obscure and sometimes rare books are entering AI-sourcing channels, but no verified public count shows how many destructively scanned copies were genuinely rare versus merely old or out of print.
Did the court say destructive scanning was legal?
For lawfully purchased print books, this court said yes under fair use. The June 23, 2025 ruling treated Anthropic’s scanning and conversion of purchased books into searchable training files as fair use. It did not bless piracy-based acquisition from shadow libraries.
Why would AI firms want physical books instead of internet text?
Books offer long-form, edited, mostly pre-LLM text that can be cleaner than the public web. That matters as labs run into AI training data limits and look for higher-quality sources beyond scraped internet content.
Further Reading
- [ORDER ON 122 FAIR USE (Bartz v. Anthropic, filed June 23, 2025)(https://cases.justia.com/federal/district-courts/california/candce/3%3A2024cv05417/434709/231/0.pdf), Primary court order describing Anthropic’s purchase of millions of print books, destructive scanning process, and fair-use ruling on purchased print copies.
- [ORDER ON 125 CLASS CERTIFICATION (Bartz v. Anthropic, filed July 17, 2025)(https://cases.justia.com/federal/district-courts/california/candce/3%3A2024cv05417/434709/244/0.pdf), Primary court order detailing Anthropic’s earlier piracy-based library building and its shift in 2024 to purchasing and scanning books.
- Anthropic ‘destructively’ scanned millions of books to build Claude, Investigation based on unsealed filings that names Project Panama, reports spending in the tens of millions, and outlines the scanning vendor’s proposed 500,000-to-2,000,000-book workflow.
- Rare book dealers fear tech firms are destroying obscure editions to train AI models, 2026 reporting on European booksellers receiving unusual bulk requests for obscure and often rare books they suspect are being sourced for AI training.
- Cuando el conocimiento de internet ya no es suficiente: la IA les pone el ojo a las librerías de segunda mano, Report linking secondhand-book purchases, intermediaries such as Zoom Books and PrepFort, and AI firms’ search for cleaner, pre-LLM-era text.
- Canadian Company Zoom Books Is Buying Out Of Print Books For AI Scraping, Trade reporting that Zoom Books was bulk-buying out-of-print books for LLM scraping.
- Feasibility study paves the way for a European Books Data Commons, National Library of the Netherlands study noting that Anthropic’s litigation records offer an unusually direct window into the current book-data market for AI training.
- Statement of Shared Practice: AI Training Requests and Unique Cultural Collections, Library-sector statement showing institutions are now formally responding to AI training requests involving rare and special collections.
- KI-Firmen kaufen offenbar Antiquariate leer – Buchhändler alarmiert, German public-broadcaster report on antiquarian booksellers seeing unusual orders routed through Zoom Books and shipped via Canada.
References
- Judge William Alsup, 2025, ORDER ON 122 FAIR USE
- Judge William Alsup, 2025, ORDER ON 125 CLASS CERTIFICATION
- Washington Post, 2026, Anthropic ‘destructively’ scanned millions of books to build Claude
- National Library of the Netherlands, 2026, Feasibility study paves the way for a European Books Data Commons
Last reviewed: 2026-07
