AI companies acquire out-of-print books for training data from secondhand bookstores

Technology companies are sourcing training data beyond public internet content, raising questions about copyright and the boundaries of fair use in AI development.

Abstract representation of books and data flows
AI-generated illustration · Sylvaris

Bookstores report bulk purchases

Secondhand booksellers worldwide have observed technology companies making large-scale purchases of out-of-print books. The pattern suggests AI developers are seeking training material not readily available through digital channels.

These acquisitions target books no longer in active publication, potentially circumventing some copyright protections that apply to currently distributed works. The purchases represent a shift from scraping publicly available internet content to acquiring physical media.

Training data scarcity drives expansion

As AI models grow larger, companies face diminishing returns from internet-sourced text. Out-of-print books offer high-quality written content that may not exist in digitized form.

The strategy highlights ongoing tensions between AI development needs and intellectual property frameworks designed before large language models existed. Authors and publishers of out-of-print works may lack awareness their content is being used for model training.

Legal questions remain unresolved

Copyright law treats out-of-print books differently than active publications, but whether purchasing physical copies grants rights to use content for AI training remains legally contested. Multiple lawsuits are testing these boundaries in various jurisdictions.

The practice adds another dimension to ongoing debates about fair use, transformative works, and whether AI training constitutes copyright infringement. Regulatory frameworks have yet to catch up with these use cases.

sources
more in Artificial Intelligence
Text-to-SQL benchmarks fail to address real-world data store complexities AI code generation tools struggle with messy production databases that lack the clean schemas found in test environments. Meta launches Content Seal watermarking system for AI-generated content detection Meta's new invisible watermarking technology addresses platform accountability for AI-generated content, though it remains less accessible than Google's existing SynthID solution. MCP servers fail agent usability testing, one-third score D or F grades Poor server design undermines the Model Context Protocol's promise to standardize AI agent tool access, creating friction in production deployments.