AI Companies Are Buying Bookshops Rarest Stock, Scanning It, and Destroying It. The Booksellers Just Found Out

Awais Khalid

August 15, 2026

AI Companies Buying Rare Books

The bookseller in Galway received the order for 5,000 titles and called it ‘bananas.’ The co-owner of Barter Books in Alnwick, Northumberland, looked at three months of unusual bulk orders for combinations of books that made no sense β€” not collected by theme, not grouped by subject, not following any pattern a human reader or a library would recognise β€” and concluded the only explanation was ‘masses and masses of AI money being squandered.’ In Zurich, a bookseller received requests for thousands of obscure titles from anonymous email addresses that stopped responding when asked who they were. The story these accounts tell, separately and together, is that the world’s used bookshops have been discovered as a data source by an industry that needs what the internet can no longer reliably provide: text that was definitely written by humans, definitely before AI-generated content contaminated the training corpus.

The Guardian reported on August 15 that secondhand booksellers across the UK and Ireland are experiencing a flurry of bulk orders from mystery buyers β€” orders that follow no recognisable pattern of human book collecting but that replicate reports from booksellers in the United States, Australia, continental Europe, and beyond. The timing is not coincidental. Court documents unsealed in January 2026, in the context of the Bartz v. Anthropic copyright lawsuit, confirmed that Anthropic had been doing exactly what the booksellers suspect: buying physical books in bulk, cutting off their bindings, scanning every page, and destroying the originals.

Key Developments

  • πŸ“š Secondhand booksellers across the UK, Ireland, and Europe are reporting unusual bulk orders, including a request for 5,000 obscure titles from a single Galway shop, which they suspect may be linked to AI companies seeking physical books as training data.
  • βš–οΈ Court documents from the Bartz v. Anthropic lawsuit, unsealed in January 2026, confirmed Anthropic ran “Project Panama” to buy physical books in bulk, remove their bindings, scan the pages, and destroy the originals. A US federal judge ruled this legally protected as fair use.
  • πŸ€– Books published before 2022 are now being commercially marketed by suppliers as “the world’s best AI training data” because they are less likely to contain AI-generated text, making newer web data potentially less desirable for some training uses.
  • πŸ›οΈ Anthropic said: “None of our data acquisition programs buy and destroy rare or antiquarian books,” distinguishing rare or valuable collector items from harder-to-find professional, academic, and STEM books, which it does not deny acquiring.

What the Court Documents Revealed

Project Panama

The scale and deliberateness of Anthropic’s physical book acquisition programme is documented in the court record. Internal planning documents, made public in the Bartz v. Anthropic case, described an effort internally called Project Panama. The documents characterised it as an effort to ‘destructively scan all the books in the world.’ The same documents indicate the company ‘didn’t want the work known publicly.’ To execute the project at scale, Anthropic hired Tom Turvey β€” who had previously worked on Google Books, the project that scanned millions of physical volumes in partnership with major libraries β€” specifically to obtain books at scale, according to the court record and The Guardian’s reporting. The mechanics were consistent across thousands of titles: physical books purchased, spines cut, pages scanned, originals destroyed. Unlike a library’s digitisation programme, which typically preserves the physical copy, Anthropic’s approach was designed to convert physical copies to digital files with no surviving original.

The Fair Use Ruling

A US federal judge ruled in 2025 that Anthropic’s destruction of the books it purchased was legally protected under the fair use doctrine. The reasoning turned on the nature of the transformation: Anthropic legally bought each physical copy, converted it into a digital file for a transformative purpose β€” training an AI model β€” and destroyed the original rather than creating an additional copy that a copyright holder could claim as unauthorised reproduction. The judge found this process analogous to prior fair use rulings on transformative use of copyrighted material. The ruling does not address the cultural or ecological cost of destroying physical copies, only their legal status under US copyright law. Separately, Anthropic reached a $1.5 billion settlement over claims relating to its earlier use of pirated digital books β€” a settlement that covers a different data acquisition method but illustrates the legal pressure that physical book acquisition was partly designed to sidestep.

Why Physical Books Are Now the Premium AI Training Asset

The AI Text Contamination Problem

The demand for pre-2022 physical books reflects a specific technical problem in AI training data that has become acute as AI-generated text has proliferated across the internet. Large language models are trained on text corpora. The quality of what they learn is heavily influenced by the quality of the training text. In 2024 and 2025, researchers began documenting a phenomenon sometimes called ‘model collapse’: the degradation in model quality that occurs when models are trained on text that was itself generated by AI models, rather than text produced by humans. The specific mechanism is that AI-generated text amplifies certain statistical patterns β€” stylistic tics, vocabulary distributions, argumentative structures β€” while compressing the distributional diversity of human writing. Training a model on a corpus contaminated with AI text risks producing a model that is, in measurable ways, less capable than a model trained exclusively on human-authored text.

Pre-2022 Books as ‘Clean’ Data

Physical books published before 2022 offer a solution to that contamination problem because they were, by definition, written before the current generation of AI language models existed. ISBNdb, one company now offering bulk book-buying services to AI developers, has marketed pre-2022 printed books as ‘the world’s best AI training data’ precisely because they are ‘structurally guaranteed’ to be free of AI-generated content, and because β€” unlike web text β€” they represent the edited, curated, and often expert-authored text that populated the corpus on which the best-performing language models were initially trained. The competitive logic is clear: if web text is increasingly contaminated with AI-generated content, and if copyright lawsuits have restricted the ability of AI labs to use digital libraries without licensing, physical books β€” particularly older, less common, non-digitised works β€” represent a relatively untapped, legally purchasable, and contamination-free training corpus.

What Booksellers Are Experiencing

The Pattern of Orders

Stuart Manley, co-owner of Barter Books in Alnwick, one of England’s largest secondhand bookshops, described orders that began arriving approximately three months ago as unusual because they did not follow any recognisable collecting pattern. Traditional bulk buyers β€” libraries, resellers, institutional acquisitions teams β€” typically order books grouped by subject, author, period, or genre. The new bulk orders contain combinations that ‘definitely strange,’ in Manley’s description: unrelated subjects, random periods, no obvious thematic logic. That pattern is consistent with algorithmic bulk purchasing designed to acquire the maximum volume of non-digitised human-authored text rather than building a coherent collection. The Galway bookshop order for 5,000 obscure titles, reported by the Irish Times, follows the same pattern. Booksellers in London’s rare book district have received enquiries from anonymous email addresses with no names or company details attached, asking for additional photographs of books before placing sizeable orders β€” and going silent when asked for identification. The increase has been replicated across Europe: Swiss, German, Dutch, and Australian booksellers report similar experiences, according to the NL Times and other regional coverage. As covered in our earlier reporting on how Anthropic’s licensing and copyright practices intersected with US export controls, the company’s data acquisition strategy has consistently provoked regulatory and legal friction β€” the physical book buying programme is the latest point of public controversy in that ongoing tension.

The Shipping Arbitrage

The Irish Times reported an additional nuance that explains the European dimension of the buying surge: transatlantic shipping costs for thousands of physical books are significant enough that buyers appear to have established collection points across Europe rather than shipping directly from European bookshops to US-based facilities. The suspicion among European booksellers is that purchased books are being consolidated at European collection points and shipped in bulk containers to the United States, where the fair use ruling makes it legal for a purchaser to scan and destroy a legally acquired physical copy β€” a legal protection that does not exist in the same form in European jurisdictions. German booksellers are ‘convinced,’ according to their own reporting, that large orders of old books are being used to feed large language models, and that the indirect routing through third-party service providers is specifically designed to keep AI companies’ names out of the purchase records.

Anthropic’s Response and Its Limits

Anthropic issued a statement in response to media enquiries: ‘Claude is trained on a mix of publicly available web data, commercially acquired datasets, and data we generate ourselves. Sourcing books is a widely used approach for training large language models across the AI industry. None of our data acquisition programs buy and destroy rare or antiquarian books.’ The last sentence draws a distinction that the court record complicates. Project Panama’s internal documents describe prioritising ‘less common’ books β€” a category that includes professional references, academic monographs, STEM texts, and humanities scholarship that is out of print and not available digitally. Whether ‘less common’ and ‘rare or antiquarian’ are distinct enough categories to satisfy Anthropic’s denial is a question the company has declined to answer in detail. The Scottish Herald’s editorial captured the irreversibility argument that sits beneath the semantic dispute: ‘Once a final non-digitized copy of a book is shredded, it’s gone forever.’

The Cultural Loss Argument

The cultural dimension of the physical book destruction programme occupies a different argumentative register from the copyright and fair use questions, but it may ultimately be the more durable concern. Copyright law provides a framework for determining whether a use of copyrighted material is legally permissible. It provides no framework for determining whether a physically unique copy of a text that exists in no other form should be destroyed so that its content can serve as training data for a commercial AI system. A Dutch bookseller called the practice ‘a kind of barbarism.’ Berlin bookseller Zerfass told the Irish Times he sees the consequences as ‘far more than a final death knell for booksellers’ β€” he describes the broader goal as one that would ‘fragment human thinking even further’ and ‘make people question even more their critical thinking abilities.’ A Scottish Herald editorial compared the situation to the burning of the Library of Alexandria and the book bonfires at the centre of Fahrenheit 451. Whether or not those comparisons are proportionate, they reflect a genuine irreversibility that is distinct from the legal question of whether the destruction is permitted.

What Retailers and Alliances Are Seeking

Independent bookseller alliances in the UK and Ireland have begun urging the development of clearer commercial guidelines to prevent automated bulk ordering systems from depleting physical inventory that is intended for general readers and collectors. The specific ask is not a prohibition on AI companies buying books β€” which courts have ruled they are entitled to do β€” but transparency: clarity about who is buying, at what scale, and for what purpose, so that booksellers can make informed decisions about whether and how to facilitate bulk purchases. The regulatory dimension connects to the EU AI Act’s transparency obligations that took effect August 2, 2026, which do not directly regulate book acquisition but which establish a principle β€” that AI deployers and providers must be transparent about their systems’ inputs and processes β€” that advocacy groups are beginning to argue should extend to training data acquisition practices. The copyright framework governing AI training data is also under active legislative review in the UK, Ireland, and across the EU, following similar controversies involving web-scraped text and image data. As covered in our reporting on Australia’s hardline AI copyright and data center laws, the legal and political landscape around AI training data acquisition is shifting across multiple jurisdictions simultaneously β€” and the physical book buying programme may become one of the more vivid test cases in those legislative debates, precisely because its physical, irreversible, visible nature makes it easier for non-specialist audiences to understand than abstract arguments about web scraping and API access.

What Happens Next

The short-term commercial consequence for secondhand booksellers is already visible: prices for out-of-print non-fiction, academic references, STEM texts, and professional monographs are rising as demand from opaque bulk buyers competes with demand from human readers and collectors. The cultural institution implications are longer-term: if the most comprehensive buyers of the physical books that populate the world’s used bookshops are AI companies acquiring training data rather than readers acquiring knowledge, the secondhand book market’s function as a circulation system for human intellectual culture is being supplemented by a new function as a raw material source for machine learning. Whether those two functions can coexist without the latter cannibalising the former is the question that bookseller alliances, cultural heritage advocates, and eventually legislators will need to answer β€” probably before AI companies have finished acquiring whatever inventory they need.

Why It Matters

The UK and Ireland bookseller story matters because it makes visible a training data acquisition practice that AI companies have preferred to conduct quietly, without public attention, and that the legal system has ruled permissible without addressing whether permissible and appropriate are the same thing. The US fair use ruling answers one question: can Anthropic legally buy a physical book, scan it, and destroy it? Yes. It leaves open a different question: should the world’s physical book collections β€” including the final physical copies of works that exist nowhere else β€” be treated as raw material for AI training, available to whoever can afford to buy them in bulk? That question is cultural, ethical, and ultimately political rather than legal, and the booksellers of Alnwick, Galway, Zurich, and Berlin are discovering that they are the first people required to answer it in practice.

Sources

The Guardian, August 15, 2026 (UK and Ireland booksellers report). Irish Times, August 10, 2026 (Galway 5,000-title order). City AM, July/August 2026 (London rare bookshops). 404 Media (original bulk-buying reporting). Snopes fact-check on Anthropic book destruction claims. Startup Fortune / Futurism analysis. Moneywise / Yahoo Finance. Bartz v. Anthropic court documents (January 2026 unsealing). Scottish Herald editorial.

Stay Ahead of AI

Get the latest AI news delivered to your inbox.

We don’t spam! Read our privacy policy for more info.