AI developers are acquiring and destroying millions of physical books to scan them into training datasets, reshaping the used-book market and inflaming copyright disputes. Intermediaries quietly source large lots—particularly pre-2023 titles prized for human-authored text—while booksellers warn rare and out-of-print works may be lost. Courts have found that scanning legally purchased books can qualify as transformative fair use, including in litigation involving Anthropic, OpenAI, and Meta. Yet a separate ruling approved a $1.5 billion settlement requiring Anthropic to pay roughly $3,000 per title for works obtained from pirated sources. The backlash has prompted some industry voices, including Elon Musk, to urge non-destructive preservation of rare volumes even as the race for high-quality training data intensifies.
Related articles:
U.S. Copyright Office: Artificial Intelligence Initiative
Authors Guild, Inc. v. HathiTrust





























