LLM & Text Generation12 min reading time

AI companies are turning old books into training data, Fahrenheit 451-style

Mashable
Read full post
AI companies are increasingly sourcing physical used books, including rare and out-of-print titles, to create training data for large language models. ISBNdb briefly offered a service to supply up to one million physical books per order for AI training, promising confidentiality, but later discontinued it. Anthropic has reportedly purchased and scanned millions of physical books, raising questions about the prevalence of destructive scanning and the fate of original books.

More on this story


More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model

Unite.AI