Sources
Stay ahead of AI art
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Discuss this with
Pick a companion and get their take on this story

Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Pick a companion and get their take on this story
Rare booksellers across the industry are reporting a troubling pattern: bulk orders of out-of-print and hard-to-find books, followed by evidence the volumes were destroyed — fueling suspicion that AI companies are acquiring physical texts specifically to scan them for training data.
According to Ars Technica's reporting, dealers describe buyers — often working through agents or shell accounts — ordering large quantities of specific, obscure titles that share one common trait: they are not available in digitized form anywhere. The orders are placed without the typical collector's interest in condition, provenance, or resale value. Books arrive, and the trail goes cold. Some sellers have later found their titles listed nowhere — not on resale markets, not in library acquisitions — suggesting the physical copies simply ceased to exist after purchase.
The legal logic, if this is what's happening, is straightforward: buying a physical book grants ownership of that object. Scanning it for personal use sits in a legal grey zone, but destroying it afterward removes the evidence of any copying. It is a far more legally defensible position than scraping copyrighted text from the web — the strategy that has landed several AI companies in high-profile litigation.
The targets appear to be books that never made it into Project Gutenberg, the Internet Archive, or any major digitization effort — niche academic texts, regional histories, specialist technical manuals, obscure fiction. These are exactly the gaps in existing training corpora that AI developers have publicly acknowledged wanting to fill. Large language models and image-generation systems trained on broader, rarer text distributions tend to perform better on long-tail prompts and specialized knowledge — the kind of nuanced output that creators using AI image generation tools increasingly demand.
For AI-art creators, the downstream effect is less direct than a model release or a pricing change, but it is real: the breadth and specificity of a model's training data shapes what it can render convincingly. A model trained on rare design history, obscure mythology, or specialist craft literature will handle niche aesthetic prompts more reliably than one trained only on what was already online.
The rare-book trade is small and relationship-driven, which means word travels fast. Dealers are comparing notes, and some have begun adding anti-scanning clauses to sale agreements or simply refusing orders that fit the profile — large quantities, indifference to condition, no stated collecting purpose. Others are alerting professional associations.
The difficulty is attribution. Buyers rarely identify themselves, and the use of intermediaries makes it nearly impossible to name a specific AI company. No major AI developer has confirmed the practice. The suspicion remains pattern-based: consistent enough across independent sellers to be taken seriously, but without the paper trail needed for legal action.
This sits alongside a broader set of questions about what AI companies are doing to assemble training data — questions that have already produced lawsuits over web scraping, disputes over the use of copyrighted images, and ongoing debates about provenance metadata standards like C2PA watermarking. Physical books, it turns out, are not outside that contested perimeter.
If the practice is confirmed at scale, it would represent a significant escalation: not just copying existing digital content, but actively converting and eliminating physical cultural artifacts to do it. The rare-book community is small enough that even a modest acquisition campaign by a well-funded AI lab could cause measurable damage to the availability of irreplaceable texts.