Books in .txt format for AI training purposes
twitter.com
twitter.com
[1]. https://www.gutenberg.org/
Fun trivia: I once took an NLP class and discovered that I could predict whether a text was written by Edgar Allan Poe or H. G. Wells with ~86% accuracy based on the presence of one word: "whereupon". The language had changed enough in 60 years such that that single word was no longer used.
British novels written in the 1950s consistently use “presently” to indicate a short passage of time. In the following decades it seems to have been removed from editors’ style guides and replaced mostly by plain old “soon”.
Are you objecting to the use of data for training, or to the compilation of training datasets?
The GP is objecting to the casualness with which privacy rights and intellectual property rights are ignored by so many in the AI community, not to a choice of whether they object to one or another specific manifestation of how the community is doing so.
[edit added]: Those who would argue "no, there is no casual ignoring of rights", consider for example the official DMCA takedown instructions for the collection [0].
Or is the trained model a derived work of those books?
Perhaps we should figure that before we start a dystopian nightmare of written material.
Training a language model, I would argue, is equivalent (in terms of being derivative) to generating a list of the top ten words in a corpus of text. Not really a derived work.
Edit: https://the-eye.eu/public/Books/humble_books_20180509/
Cue the "they're depriving me of my income" complaints from authors.
That was not the question.
However, the legal environment is very different than in the commercial world or consumer piracy, as it generally involves various legal exemptions (differing between locales) that do allow such usage. For example, I work in NLP research on aspects that involve handling large corpora of copyrighted text. It's easier to do it with cooperation of the publishers for various practical reasons, however, we still can and do use also the works of the publishers who would refuse to grant any permission, because local copyright law has specific exceptions that allows the usage these works for noncommercial research purposes. Doing so is not ignoring their rights, their rights are not violated but rather they are limited; their exclusivity right (privilege would be a more appropriate word) to make copies is not absolute. There are even some countries with explicit legal duty for the publishers to provide digital versions of their works to national corpora where they will be used for (among other things) machine learning models.
The specific consequence, however, is that we can't legally share the full datasets which we are using with the public, like it was done in the original post with this particular dataset, as that would be a violation of the publishers' commercial rights; we can provide them to specific researchers for limited noncommercial purposes only. But I can download this dataset or one like this and use it my research legally; just as I can rip up a physical book, scan it, make a digital copy and OCR it, and use it in a research corpus (with copies distributed to other researchers) even if the publisher disapproves.
It'll work well until we can upload ourselves to the cloud and we'll have to revisit it
It's a grains-to-heap problem; eventually, you have to make an ugly choice.
But yeah that’s separate from the question of whether you properly licensed the data to train on in the first place.
A big chunk of the computing community seems to approach licensing as “I can see it, so I can use it.” (See GPL code used where it shouldn’t be.)
You're not generally distributing, performing, &c. your state of mind. On the other hand, if you then go and generate work based on what you read, then distributing that work certainly can be infringing. Thus "clean room" techniques, where one team reads copyrighted text, writes up a spec (which may then be checked off by lawyers), and then another team, without reading the copyrighted text, implements based off the spec, are sometimes used to attempt to launder copyright taint.
Taking a step back, the fact that well-known authors have infamously declared that they don't read fanfiction as a CYA move against accusations of plagiarism suggests that yes, brains that have read text are widely considered derivative.
I don't think I agree. My employer values my contribution in some part as an oracle. People at work ask me questions and I answer them. Those answers come from the sum of my experiences (a biochemical 'model'). Other people more directly conduct public performances of their talent.
If nothing else this will likely blur the lines on what's considered original work.
Hell, unless you paid for it, you shouldn’t even have the collection to train on, whether you distribute it or not.
Effectively, everyone’s focusing on whether your story(/model) is sufficiently transformative from Die Hard to be fair use for you to distribute it, without addressing the complaint that you snuck into the theater too, and held the door for others to follow me in.
It's not that simple unfortunately, and companies could still send a C&D/threaten invoking the CFAA, and it would still be a mess. Although in practice it's not worth the effort for companies to sue as long as the scraper is not monetarily benefiting from it (which is how the court case happened).
GPT-3 was trained on the Common Crawl + books.
With that said, does anyone know if the filename structure inside that file makes any sense for a human? Can you find a book by author/title?
Also, what's in that dmca.mp4 file? All I hear is music.
Presumably this is a "fuck you" to the DMCA law.
EDIT: there is unfortunately no "suitable finale."
Also the dmca.mp4 is a video of like 5 or 10 people singing into microphones while pretending to jack off
I wish you could see it, here's a basic description.
It begins with a close up of a person singing while performing a gesture of masturbation then zooms out to three people signing and then a full group of around 10 people. It's basically telling whoever is trying to DMCA to jerk themselves off.
I spent four or five days intensely working on the script to convert an .epub to .txt. It's a surprisingly hard problem, so I tried to be correct in every detail. It was important to me that an AI learn the way humans learn.
Therefore, yes, I have found that these books are quite readable. However there are some unique challenges due to the nature of it being .txt, which I would love to help you navigate.
Firstly, the top and the bottom of each file is usually "repetitive". By that I mean, there tends to be a lot of metadata that might be hard to skip.
However, the neat thing is, these aren't actually .txt files -- it's Markdown. And we can take advantage of the markdown format to help you out.
Here is a direct link to "A Mathematician's Apology": https://gist.githubusercontent.com/shawwn/1b399325e866731165...
I like this file as an example, because it illustrates both problems that will be challenging for you. Firstly, the book itself doesn't begin until halfway down the file. There is a ridiculously long "forward" section, written by someone else. Normally this would be easy to skip.
The cheat code I use to read this book is to search for "# 1". Double quote, followed by pound sign (or "hashtag" as the kids say nowadays, ha), followed by a space, then the number one.
That brings you straight to chapter one.
However, not all books use that format for chapter one. So I would recommend searching for a regex, if you can: "beginning of line" followed by pound sign.
That'll usually do the trick. And yes, I find all of these to be quite readable, which I was proud of.
The code is courtesy of Aaron Swartz, by the way. I merely refined it. I am so proud of him for what he was able to achieve during his lifetime. This is "html-to-text" with some modifications.
If you want to convert .epub files to .txt for your own use, the script is here: https://github.com/shawwn/scrap/blob/master/epub2txt-all
It has a rather cryptic name of "epub2txt-all". Sorry about that. But the script itself should be just a matter of running it.
Here are some notes on what I did with this script, in case you find it helpful: https://github.com/soskek/bookcorpus/issues/27
Happy reading! Please let me know any other questions you might have.
By the way, dmca.mp4 is a joke video. Sort of. It's a good question of what is actually going on there. As far as I can tell, it's a dozen professionals who have gathered together to simulate jacking themselves off while singing "ahhhhhhHHhhH" for ten minutes. I don't really know anything beyond that, but the-eye.eu community seems rather proud of it.
- 0/0.4 - Mike Lancaster.epub.txt
- 0/0.721 - Gary Webster.epub.txt
- 0/01 - Alec Dunn.epub.txt
- 0/01 Kai_ Ninja Of Fire (Scholastic) - Greg Farshtey (retail).epub.txt
- 0/02 Crescendo - Becca Fitzpatrick.epub.txt
- 0/03 Cole_ Ninja Of Earth (Scholastic) - Greg Farshtey (retail).epub.txt
- 0/03 - Jean-Christophe Valtat.epub.txt
- 0/04. R. W. Peake - Antony and Cleopatra Part I Antony (Marching With Caesar, Book 4) [Retail].epub.txt
- 0/05. R. W. Peake - Antony and Cleopatra Part II Cleopatra (Marching With Caesar, Book 5) [Retail].epub.txt
- 0/05 LEGO Ninjago - Snake Attack! (Scholastic) - Tracey West (retail).epub.txt
Apologies, I had included the wrong corpus originally, the correct list is this one.Second edit - Books3 is much cleaner than Books1, I am editing my negative opinion to a more positive one to reflect that.
However, saying: "Suppose you wanted to train a world-class GPT model, just like OpenAI. How? You have no data. Now you do. Now everyone does"
Is quite a stretch. You need much, much more that 36GB of quality data to train a world-class GPT model.
Openai states themselves that they trained on "40GB of internet text"
Your link is for GPT-2, GPT-3 used much, much more data.
GPT-3 was trained in part on "books2" corpus which is not public but seems to basically be the same thing as this: 200k books * 100k words per book on average * ~3 token per words = 60B tokens, books2 is 55B tokens so it checks out.
The total amount of tokens that GPT-3 was trained on from all sources is a combined 500B tokens, this is merely 10% of what they have.
Obviously our wetware is optimized for it and we aren't true blank slates, but the enormous magnitude of the discrepancy makes me think we're on the wrong path.
Instead, as you pointed out, it's just proof of how efficient our wetware is.
The ironic thing is that ML became popular because we didn't have the technology back then to properly model actual neuronal activity. Now we do, but ML is so established that modeling "true" intelligence is no longer de jure.
https://www.technologyreview.com/2020/08/22/1007539/gpt3-ope...
So, GPT? I think that's probably a form of intelligence too, but not one that's particularly similar to our own. I also doubt that we'll have much luck getting it to ever perform at near-human levels, when you consider the data and power requirements. I think it's not simply a matter of silicon vs wetware; I think the GPT approach is substantially more different from our own than the hardware it runs on. It's a form of intelligence but that doesn't mean it is like our own.
That's not a fair comparison. The human neural network is not trained _from scratch_. There is a very good fundamental structure we inherited through millions of years of evolution.
It is worth noting that the training data used for training humans has been undergoing it's own optimization process for a while too (arguably for about 40ky, since the Upper Paleolithic Revolution, which is roughly when cultural evolution started taking over). In AI adjusting the training data is called Curriculum Learning. Right now most curriculum learning is done just by adjusting the order of training samples, rather than creating samples specifically optimized to facilitate learning (although GANs might be considered a Socratic approach to the latter, if you squint at it).
books3 may be 10%, but The Pile is building the rest:
https://twitter.com/arankomatsuzaki/status/13204141418954874...
https://github.com/EleutherAI/The-Pile
https://www.eleuther.ai/get-involved
https://media.discordapp.net/attachments/735217892517216366/...
https://media.discordapp.net/attachments/735217892517216366/...
I just have a minor gripe with saying that "now we can train world class GPT model" thanks to that, as you said it's just one piece, and as a typicial HNist I had to point it out :).
But after spending roughly one year acquiring knowledge related to this work, I feel I can say with a fairly high degree of certainty that this dataset alone is enough to train a model that will achieve "world class" status in some area. Writing books, perhaps.
Which part of my logic do you feel is mistaken, and why? I am actually quite interested to hear thoughts from someone who is very pedantic about such things.
Time will tell if fear would have been wise. But for now, I'm curious to see whether "the link will last for years" is true.
If it goes down, I'll append instructions to the original twitter thread on how to access it elsewhere. I also don't have any experience setting up the alternatives you mention; if you do, please feel free to mention it somewhere (twitter DM is always a reliable way to reach me) and I'll highlight it.
As for a torrent link specifically, I did request that the-eye have a torrent ready to go on day one, for just such an eventuality. I was confused when they didn't seem worried. After spending almost an hour describing in detail the kinds of repercussions that might inevitably follow, he simply said that fear does not control him the way it controls me, and that he will ensure it remains. (Probably via torrent, if such a thing becomes necessary.)
I've thought a lot about what he said, and how he phrased it. I decided to trust them to take care of the data. Mostly, though, I made the decision out of intellectual curiosity to see how true their ambition really is.
In the meantime, joining their community on Discord and/or donating to them would be helpful (though to my surprise, they didn't seem too interested in monetary concerns either). http://the-eye.eu/
I've developed a web platform that handles books online in a very unique way, and my demo 'book' for that technology was War and Peace:
https://quanta.wiki/n/war-and-peace
but probably if I need to add more books I can go to Project Gutenberg. Anyway, appreciated the reply
AFAIK, most is in PDF, EPUB or mobi format. I'd presume it's not too difficult to extract text from the latter two, but extracting text from PDFs is far from simple, and something you get working for 1 PDF won't necessarily work on another.
Also, a lot of books that were behind the bibliotik signup wall have been freed by /u/-Archivist.
https://www.reddit.com/r/opendirectories/comments/f2teym/pro...
I personally recommend you start with MyAnonamouse.org. They hold invite interviews twice a week over IRC and I have heard that the community is really welcoming to new users.
I would hope for UTF-8, but given the old-school, Windows-y .txt extension, it could just as easily be Windows-1252 or something. LF or CRLF line endings?