A closer look at BookCorpus, a key dataset in machine learning
towardsdatascience.com
towardsdatascience.com
This isn't to suggest that organizations shouldn't invest in actively understanding their training data, but that post-hoc bias analysis is going to be a critical component of evaluation for the foreseeable future.
What's the value of these scant few thousand unpublished romance and fantasy novels in the context of the rest of the corpus -- vast scrapings, all of Wikipedia, etc.? A sample of how people write? Why aren't more public domain works included?
The Pile (the 800GB dataset by Eluther AI) contains BookCorpus2, along with two much larger datasets of books (and a whole lot of not-book stuff). From their paper [0] the reasoning for the book datasets is that they are "are invaluable for long-range context modeling research and coherent storytelling". The reasoning for including BookCorpus2 next to Books3 and Project Gutenberg boils down to "no significant overlap with the other datasets, and others use it".
In general books are a great source of extremely high quality long-form content. They are longer than most content found on the web, and are generally of high quality, having gone through many revision rounds between editor and author. Just that both of these aren't really true of BookCorpus. Even a dump of highly rated stories from fanfiction.org might be better.
For example, the total number of tokens contained in the physical and digital collections of a moderately-sized university library is (probably) equal to or on par with the size of the training data for GPT 3.5.
What would happen if you could train just on that? I know we're using huge training sets, but how much of it is just junk from the internet?
(There should be some representative junk in the dataset, but nowhere near the majority.)
The cynical me is thinking that because this is inconvenient news and would require rework, a lot of people would prefer to suppress or ignore the author's findings (assuming true).
I note neither this paper nor any discussion of "BookCorpus" or even "book corpus" has appeared on HN previously.
Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus, 2021, Jack Bandy and Nicholas Vincent
In this case, the modified title utilizes the more descriptive language of the article's subtitle. Editors edit.