What's Actually Being Lost When AI Destroys Books — And What Isn't
Somewhere between a Reddit thread about mysterious bulk book orders and Elon Musk telling his own engineers to scan rare volumes "the hard way," a genuine story about how AI companies source training data turned into something closer to a moral panic. Headlines compared Silicon Valley to the Nazis. Michael Burry called it "evil incarnate." Substack essayists reached for Fahrenheit 451. And underneath all of it sat a real, well documented practice, "destructive scanning," that is both less sinister and more quietly troubling than the loudest version of the story suggests.
This is an attempt to hold both things at once. Take seriously, as anyone who cares about the survival of physical books should, what's actually happening to them in the age of AI. And be honest about how much of the current outrage is running ahead of the evidence. Getting the facts wrong doesn't help the case for protecting books. It just gives the industry an easy exit, a shrug and a "well, that specific claim wasn't true," followed by silence on everything that was.
What we actually know
The factual core of this story comes from a real court case, Bartz v. Anthropic PBC, decided in a northern California district court in late July 2026. Unsealed exhibits revealed that Anthropic ran an internal effort, candidly codenamed "Project Panama," to buy physical books in bulk, strip their bindings, scan the pages, and pulp what remained. An internal memo was blunt about the secrecy; the company didn't want the project's existence known. Judge William Alsup ultimately ruled that this practice, buying a book, digitizing it, and discarding the physical copy, counted as "transformative" fair use under US copyright law, and so was not itself infringing. The much larger sum Anthropic ended up paying, a $1.5 billion settlement, was for something different, training on pirated ebooks obtained before the company pivoted to buying physical copies instead.
That distinction matters and gets flattened constantly in the coverage. The piracy was the legal problem. The destructive scanning of legally purchased books was not found to be a legal problem at all; a judge with no particular incentive to be soft on Anthropic said so explicitly. But it's worth being clear about what that ruling actually settled. A court decided the practice doesn't violate copyright law. It said nothing about whether industrially pulping books after stripping them for parts is a good thing to do to the physical record of what humans have written, and "a judge allowed it" is doing a lot of unearned work whenever it's used to end that conversation rather than start it.
Why destroy the books at all, rather than scan non-destructively or license ebooks? Because slicing off the spine and feeding loose pages through a high-speed scanner is faster and cheaper than either negotiating licensing terms with publishers or handling a book carefully enough to keep it intact. That's the whole justification, throughput. Non-destructive scanning equipment exists and has existed for years; it's just slower and it costs more, and at the volumes AI companies are operating, that difference in unit cost is apparently worth an irreversible loss. It's worth sitting with how thin that trade-off actually is. This is a permanent, one-way decision made almost entirely on the basis of scanning speed.
Where the panic outran the facts
The viral version of this story escalated well past what's actually documented, and it's worth being precise about where.
The trigger was a 404 Media report describing unusual bulk-buying patterns at secondhand bookshops and naming ISBNdb, a book database company, as a possible buyer acting on behalf of AI firms. As The Atlantic's Alex Reisner dug into it, the picture got considerably murkier. ISBNdb told Reisner it had never purchased, scanned, or destroyed a single book, and that a page briefly advertising bulk-buying services for AI developers had been testing demand for a service that was never actually launched. No bookseller Reisner spoke to had records of an order from ISBNdb.
The company that does show up repeatedly in booksellers' accounts, in the US, Canada, Germany, New Zealand, and Australia, is a Canadian outfit called Zoom Books, which markets itself as a book recycler operating at industrial scale. Zoom has denied any involvement in digitizing or destroying books, saying it acquires and resells books intact, recycling only what can't be rehomed. Whether Zoom is quietly supplying AI companies is unclear; the company cites confidentiality agreements and won't say who its customers are. A logistics firm called PrepFort, which some of Zoom's shipments have been routed through, gave the same non-answer. Anthropic, for its part, told Guardian Australia it has never bought from Zoom Books and that none of its sourcing programs involve rare or antiquarian material.
So the honest state of the evidence, as of this writing, is this. We know Anthropic destructively scanned books it bought in 2024. We do not have solid evidence linking the current wave of unusual bulk orders to any specific AI company, and the "millions of rare books being destroyed" framing that dominated social media doesn't hold up well against what booksellers are actually describing. Charlie Becker, a Houston bookseller who fielded a run of large orders, said the books involved were overwhelmingly ordinary nonfiction from the 1970s through 90s, an outdated Denver travel guide, a manual for a defunct word processor. Cheap, plentiful, unglamorous books, not first editions or the last surviving copy of anything. One seller in Australia noted a $9 poetry anthology that had sat unsold on the shelf for twenty years going out in a bulk order alongside a soil-mechanics manual and a local history of a Melbourne suburb, hardly the profile of rare-book plunder. There are exceptions worth flagging honestly, too. A Singapore-based firm called 2077AI reportedly sought out niche academic titles with small print runs, the kind of specialized, low-circulation books that could reasonably be called rare in a meaningful sense. That's a real edge case, not the dominant pattern.
None of this means nothing troubling is happening. It means the specific claim driving the loudest headlines, that AI companies are hunting down and destroying the world's irreplaceable rare books, is currently unproven and, based on the buying patterns booksellers describe, probably overstated. Fear-mongering doesn't require lying; it just requires taking the most alarming possible reading of ambiguous facts and presenting it as settled.
The book-burning comparison, and where it breaks down
The instinct to reach for history, Qin Shi Huang's burning of rival philosophical texts in 213 BCE, the burning of the Library of Alexandria, the Nazi bonfires of 1933, ISIS destroying manuscripts in Timbuktu, is understandable. Book destruction has a long and ugly pedigree, and Richard Ovenden's Burning the Books makes the case persuasively that libraries have always been targets precisely because they underwrite collective memory, legal legitimacy, and cultural identity. Destroying them has historically been an act of power, an attempt to make certain ideas, or certain people's claim to a shared past, simply cease to exist.
That is not, on the evidence available, the same act as historical biblioclasm, and it's worth being precise about the difference rather than collapsing the two. Historical book burning targeted content, ideas an authority wanted gone from circulation entirely. What Anthropic did runs the opposite direction on that one axis. The content was the entire point, valuable enough to buy the physical book specifically to extract it. Nobody was trying to stop anyone from reading The Insider's Guide to Metro Denver.
But a conservator doesn't get to stop the analysis at intent, because the outcome for the object is identical either way. A unique physical artifact, with its own paper, binding, printing history, marginalia, and provenance, is gone, permanently, and nothing about "we only wanted the text" brings any of that back. Intent changes how we should judge the people doing it. It changes nothing about what's been lost. A book pulped so its ideas are eliminated and a book pulped because its ideas were valuable enough to extract and then discard both end at the same place, one fewer physical copy of that object in the world, forever. If anything, treating "we weren't trying to destroy it, destruction was just incidental to what we actually wanted" as exculpatory should worry a conservator more, not less. It means the object's survival was never even a consideration serious enough to weigh against efficiency, only a cost that lost.
The Substack argument that this makes AI companies "worse" than historical book-burners because they hid the practice is a clever line, and it doesn't hold up cleanly either way you take it. Secrecy driven by reputational risk is a different animal from secrecy meant to control what a population knows. But the more useful point buried in that argument is this. A company that instructs itself, in writing, not to let a practice become public is a company that already suspected the practice couldn't survive daylight. That instinct was correct. It shouldn't take a leaked court exhibit to find out how a company is sourcing the material it trains its products on.
Where the historical parallel does hold up, and where the Guardian opinion piece by Yale's Kathryn James lands a real point, is on the absence of any legal or cultural framework for protecting the book as a physical, historical object once its informational content has been extracted. The US has essentially no regulatory apparatus treating a printed book as heritage in its own right, separate from its text, no equivalent of protections that exist for buildings, artworks, or archaeological sites. That gap predates AI, but AI has just given it commercial teeth for the first time at meaningful scale. A book carries provenance, marginalia, printing history, and physical evidence of how it was made and used, none of which survives a scan. Whether a "wholly human-authored, pre-2022" physical book should be treated as a preservation-worthy category the way we treat rare manuscripts is a genuinely open policy question, and it's one almost nobody was asking before this year.
What's actually worth worrying about
Strip out the exaggerated "millions of rare books" framing and there's still a real story underneath, made up of a few distinct threads.
The demand for pre-2022 text is not going away, and it's structural. AI companies specifically want books published before generative AI existed because of "model collapse," the well documented phenomenon where models trained on AI-generated text degrade in quality across generations, inheriting and amplifying each other's errors. Physical books are attractive precisely because they're a large, professionally edited corpus that's guaranteed to predate that contamination. That incentive isn't going to soften; if anything it should be expected to intensify as more of the open internet fills with synthetic text.
The opacity in the supply chain is real, whether or not any given rumor about a specific buyer checks out. Multiple companies in this chain, Zoom Books, PrepFort, whoever 2077AI's actual clients are, have declined to say who they sell to or on whose behalf they buy, citing NDAs. That's a legitimate business practice, but it also means booksellers genuinely can't tell whether a bulk order is an AI company, an arbitrage reseller, or something else, and neither can journalists trying to report on it. Ingram Content Group's decision to offer publishers an opt-out from AI training sales, while candidly admitting it can't guarantee compliance given how anonymized the buying chain has become, is a tacit admission that nobody currently has full visibility into where books bought secondhand end up.
"Not rare" is being used to mean "doesn't matter," and that's a conservation error, not a factual one. Most of the coverage, including the more careful debunking pieces, treats the discovery that these books are mostly ordinary nonfiction as basically reassuring, nothing irreplaceable, move along. That's the wrong lesson to take from it. Library and archive science has never defined heritage value solely as scarcity. A run-of-the-mill 1991 WordPerfect manual or a regional Denver travel guide is, in aggregate with millions of others like it, part of the material record of what an ordinary person was reading, buying, and being sold in a given decade, evidence of print culture, not just information to be extracted from it. Losing any single copy isn't a tragedy. Losing that category of object at industrial, sustained scale, indefinitely, with no accounting of how many copies remain anywhere once the buying is done, is exactly the kind of erosion that's invisible until it's too late to reverse, because nobody was tracking it as loss while it was ordinary.
There is no meaningful regulatory distinction, in most jurisdictions, between destroying a mass-market paperback and destroying a book that happens to be the last surviving copy of something. Germany's Publishers and Booksellers Association has taken the harder line that any scanning without explicit permission violates copyright regardless of purpose, a genuinely different legal standard from the US "transformative use" doctrine that let Anthropic off the hook. That divergence is likely to become a real point of friction as AI companies operate across jurisdictions with very different copyright philosophies.
The industry's public messaging has not been straightforward. A company explicitly instructing itself, in writing, not to let a data-sourcing project become public knowledge is not a great look, irrespective of whether the underlying practice turns out to be legal. That kind of unforced opacity is exactly what invites exaggerated readings later, because it signals the company itself thought the practice couldn't survive scrutiny, even if, in Anthropic's case, a federal judge ultimately concluded it could.
Holding the line between concern and panic
The most useful thing anyone can do with this story right now is resist the pull toward the two easiest narratives, that this is nothing to worry about because a judge said it's legal, or that it's a five-alarm cultural emergency on the scale of history's great book-burnings. Neither is true, and neither is the point. The point is that a legally permitted, industrially scaled, and largely opaque practice is currently allowed to permanently destroy physical books, millions of them, by Anthropic's own account, and likely far more once every AI company doing something similar is counted, with no regulatory framework anywhere treating the book-as-object as something worth weighing against the convenience of the company doing the destroying. That would be true whether or not a single one of those books turns out to be rare.
Getting the exaggerated claims right matters precisely because the real complaint is stronger than the exaggerated one, not weaker. "They're burning the last copy of a rare medieval text" is a claim that can be checked and, so far, mostly hasn't held up. "An entire industry has quietly normalized permanently destroying physical books at industrial scale, through a supply chain designed to make it impossible to trace, with no law anywhere requiring anyone to even ask whether a given copy is the last one," is a claim that's already true, doesn't need a single unverified rumor to support it, and should be harder to wave away than a debunked headline about rare books. A conservator doesn't need the Nazi comparison, or the missing evidence about ISBNdb, to make that case. The court record alone is enough.

