AI companies acquire out-of-print books for training data from secondhand bookstores
Technology companies are sourcing training data beyond public internet content, raising questions about copyright and the boundaries of fair use in AI development.
Bookstores report bulk purchases
Secondhand booksellers worldwide have observed technology companies making large-scale purchases of out-of-print books. The pattern suggests AI developers are seeking training material not readily available through digital channels.
These acquisitions target books no longer in active publication, potentially circumventing some copyright protections that apply to currently distributed works. The purchases represent a shift from scraping publicly available internet content to acquiring physical media.
Training data scarcity drives expansion
As AI models grow larger, companies face diminishing returns from internet-sourced text. Out-of-print books offer high-quality written content that may not exist in digitized form.
The strategy highlights ongoing tensions between AI development needs and intellectual property frameworks designed before large language models existed. Authors and publishers of out-of-print works may lack awareness their content is being used for model training.
Legal questions remain unresolved
Copyright law treats out-of-print books differently than active publications, but whether purchasing physical copies grants rights to use content for AI training remains legally contested. Multiple lawsuits are testing these boundaries in various jurisdictions.
The practice adds another dimension to ongoing debates about fair use, transformative works, and whether AI training constitutes copyright infringement. Regulatory frameworks have yet to catch up with these use cases.