Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.