Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.
The goal here is to have all human knowledge in a single file, which is pretty neat IMO.
I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.
Edit: Actually, I have real evidence, the Bartz in "Bartz v Anthropic" is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.
They're scanning millions of books.
It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.
LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.
They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).
Books are more likely to be about a specific topic or story or time or setting and be more information dense
I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.
There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.
Lots and lots of information is not online. You'd be surprised.