The Battle over Books3
wired.com
wired.com
Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legislation, the AI companies would risk losing access to content from small creators without the means to sue too.
We seem to be moving into a world where corporate content cannot in any way be reused, remixed, or even archived. You cannot even own a copy - it is only accessible for a monthly fee and can disappear at any time. Meanwhile, anything created by independent creators is fair game to steal and rip off. Copyright was intended to promote and protect human creativity, but instead we got a rent-seeking mechanism used to stifle original creation.
https://janefriedman.com/i-would-rather-see-my-books-pirated...
It’s almost as if governments (even capitalist ones) work with industry and concentration of power perpetuates this kind of consolidation further. They keep us distracted so we don’t have enough collective willpower to get together and demand reform, or even better — create our own alternative open ecosystems.
I write about many other examples of government-industry distracting us here: https://magarshak.com/blog/?p=362
this is the point of a class action suit isn't it?
if it turns out training isn't fair use then Microsoft/Google/OpenAI will suddenly have class action suits for billions if not trillions of damages against them
($150,000 damages per willful infringement, after all)
the aim would be to make training on copyrighted material legally toxic and render all existing datasets and trained weights unlawful
the damages are simply a bonus
it's like patent trolling, but with no expensive patent required
"We bought these weights in good faith from Digitus Tertius corp of bermuda."
Look, I get that unethical corporations using this pirated training data for their artist-usurpation machines is bad, right? But EVERYONE being able to dismiss the rights and wishes of current artists while they work to create artist-usurpation machines of their own? That's not any better! You don't need to "democratize access" to that!
Funny, I get the same tingling feeling when I see the phrase "restrict access." I guess it's an unreliable signal, huh.
Generally speaking, democratizing access is about removing monopoly power that extracts illegitimate rents -- and that tingles in a good way.
Libraries have always been at the forefront of democratizing access, and hasn't that been an unadulterated good? Or whether it's community colleges that democratized access to higher ed, or the deregulation of air travel that democratized air travel through cheaper prices that made it more available to the masses, instead of artificially restricting routes.
I can't think of a single example of "democratizing access" that is "sketchy at best" or "outright evil". They all seem pretty great to me!
But I guess I'm also firmly on the side of AI training here -- I see no reason for additional compensation to authors/artists for training on their data, when anyone can go to a library or museum and then go and create their own works influenced by what they've seen. Who cares if I hire an expensive consultant who's read a lot of books at the library, or a cheap AI who's read a lot of books from Books3? Why would authors/artists deserve extra compensation for the latter but not the former? There's just no clear legal principle behind that.
Then there's also the issue with things like art, music, and code. Where does the line fall with scraping Github, Soundcloud, DeviantArt, or Instagram and using things like that without permission? Most of the code on Github is open source, but there's a lot of difference between the GPL and BSD licenses.
No it's not at all, except in extremely limited circumstances.
When George Lucas made Star Wars, did he cite all the Westerns and space opera serials and movies that influenced him? When you give a presentation at work on why you should move to a sharded database, do you cite the history of academic work on sharded databases? When you use Times New Roman in a document, do you cite the British newspaper The Times, or Robert Granjon's prior serif designs from the 1500's?
Of course not.
Legally, you can do whatever you want with ideas and styles and whatnot, which is what AI is about. Legally, you only run into problems when you reproduce sections of copyrighted works verbatim, without a license, in a manner that's not considered fair use. Your answer to "where does the line fall" is quite clear legally -- it's the line demarcated by fair use, which has nothing to do with licenses. AI doesn't change that.
> A “derivative work” is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted. A work consisting of editorial revisions, annotations, elaborations, or other modifications, which, as a whole, represent an original work of authorship, is a “derivative work”.
As I understand it, derivative works must be created with the legal use of the original work, or be fair use, otherwise they are infringing.
If you take a book and turn it into a movie, that's a derivative work. Anyone can see the direct resemblance -- the transformation or adaptation.
But if you take a book, convert each letter to a number, add up the numbers that make each sentence, and then sell that as a list of "random" numbers, that's not a derivative work. The end result is sufficiently transformed that copyright no longer applies. Ownership of the original work has no relevance.
And AI weights are like that. They're a complete transformation. They're not a derivate work. The only thing you have to make sure of is that they haven't been overtrained to the extent that they can regurgitate whole chapters of the texts they were trained on, for example. But that's not something they're currently able to do, and obviously copyright law will force companies to ensure it stays that way. (Not to mention that companies would do it anyways, due to the economic motivation of reducing model sizes to cut costs.)
Now, I understand the argument that perhaps the specific work has been homeopathically diluted down to nothingness in the weights and so therefore has only been used to contextualise the compression process of other works, but if the weights can be reasonably used to generate copyright infringing text (and condensations and abridgements and transformations are explicitly listed in the law, verbatim copying is not necessary), or even answer substantial questions about it, then that shows that the weights included that data.
If I take a sound file and compress it down so it's poor quality but I can still make out the tune, that doesn't mean that I've avoided copyright law.
No they're not -- they're more like the dictionary generated to produce a lossless compressed data set. But then we throw out the compressed data itself, and keep only the dictionary.
> but if the weights can be reasonably used to generate copyright infringing text (and condensations and abridgements and transformations are explicitly listed in the law, verbatim copying is not necessary)
First of all, they haven't been shown to substantially generate infringing text that aren't the kinds of short snippets covered by fair use. And my previous comment already explained that longer texts are not going to happen, for both legal and economic reasons.
But secondly, you're wrong about "condensations and abridgements and transformations". You can absolutely sell a page-long summary of a book without getting permission, for instance. What do you think things like CliffsNotes are all about? Or all those two-page "executive summaries" of popular busines books?
You can't abridge a 1,000 page book to 500 pages and sell that, but you can summarize its ideas in a page and sell that. Which is basically the approximate level of understanding that LLM's seem to absorb.
the problem with this as an example is that copyright would not apply to this transformative work, not the original author's copyright nor your new authorship because this transformative work contains no creative human expression (unless the original book was designed to add up to some fortune cookie, of course, in which case you have not transformed it)
A nuttier, chewier example would be retelling a litigious story like Moana ("consider the copyright, across all these leaves... make way!"), from the pig's perspective or something, and seeing what would fly and what wouldn't.
Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "enshittification". (The OALD 7ed in it's foreword even states "Corpus analysis now makes it possible to draw authentic examples from a vast range of attested contemporary usage. A concordance will display hundreds or thousands of them to choose from.")
I suspect that the inclusion of a few sentences -- especially those that introduce a new word or usage of a word -- are fair use, but the inclusion of the entire texts is not.
This then brings up an interesting point where the computer scientists/linguists developing tools like WordNet or other NLP databases would be at an advantage to those that take the approach of throwing a lot of data into a neural network and hoping for the best. Yes, it is a lot more work/effort to develop those NLP databases, but in the end they may end up being more robust, especially around the question of copyright.
In essence, the long established principles for text analysis is that facts about text (concordances, collocation statistics, n-gram counts) are neither copyrightable nor derived work, and thus can be calculated, gathered, used and distributed even if copyright holders of the source data object. Now a court might judge that training a large language model is substantially different or that it's effectively the same, but such a decision wouldn't affect corpus linguistics and how they use sample sentences, only whether LLMs get the same treatment or not.
They are obviously transformative (the whole "dictionary" part), copy factual information (use of a word, not what the sentence itself is saying), are not substantial (one sentence out of many thousands), and do not impact the original work's value (nobody would buy a dictionary instead of a novel because it contains a sample phrase from that novel). The OED cites its sources, which also strengthens its case.
Compare that to AI, which is more than happy to write a short story to the prompt "Write a story about Bucky and Captain America falling in love, and living happily ever after in a mountain cabin." (Transformative? Maybe. Factual? No. Substantial? Yes. Impacts value? Yes.) Works like the AI's output have been dealt with in lawsuits like Salinger v. Colting, and it simply is not allowed. The big question right now is: what about the AI model itself?
(We noticed The Pile was recently taken offline, so we hosted it: https://thenose.cc. Apparently Books3 was also a part of The Pile, so feel free to download.)
I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that a professor hosted pirated content under his own name under the guise of AI training?
This is the kind of effort that could have been done anonymously, just as all of the pirated books had already been uploaded and hosted anonymously. I’m not sure why he expected any different outcome by re-pirating everything under his own name.
The journalists seem to be loving it, though. All of the tech journals have an article about this guy.
They are essentially the same content (that is, the documentation for the copy of books3 seoarately hosted in huggingface says that it is all of Bibliotik in plaintext form, presumably as of a particular point in time.)
> This is the kind of effort that could have been done anonymously
Sure, its something each group training an AI could do independently at greater aggregate cost until someone succeeds in taking the original source down, but not only would that be costlier, but it in would involve less transparency and comparability across model architecture, or at least required the transparent, comparable trained version to be different from the full version.
I meant he could have uploaded it under a pseudonym rather than broadcasting to the world that he was the one doing the uploading.
>Surely there must be more to the story?
I asked him a while ago about this. The story is as simple as it sounds:
>Whether the defendant had purchased a signed copy or flagrantly shoplifted a dog-eared paperback wouldn’t matter during arguments over whether The Bedwetter, Too was a derivative rip-off or a transformative parody.
This strikes at the heart of why this case is about to be laughed out of court.
The argument the plaintiffs are making is that ChatGPT is a "derivative work", i.e. letting people use the software is akin to distributing carbon copies of the book at issue, with at most slight modifications (typical derivative works include translations, screenplay adaptations, etc.).
Since ChatGPT obviously cannot literally produce the full text of the book on command, the very strained position they're trying to advance is that short, several-paragraph summaries constitute a derivative work.
That is to say, they're arguing that writing, say, a review, or a book report, is an act of copyright infringement tantamount to taking a book, translating it into Japanese, and selling that translation.
It's a deeply stupid and wrongheaded argument, and it deserves to die a quick death.
The problem with my question is the following:
Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to other sites, which are unlikely to block GPTBot themselves. How will DigitalOcean (or any other company) address this?
It looks to me like it's Catch 22 in every direction you look, and unless you're someone like The New York Times who can afford to outright protect the data with licensing...
It's just something thats been on my mind lately but I don't understand the finer details of it.
Steal or do something shady, raise enough money / power so by the time your noticed, you have the money to win in the courts.
You already see the propaganda about “none of this should matter because the cure for cancer is on the way courtesy of AI.” I mean maybe it it but it smells fishy to me.
The would possibly need to apply some effort to appear human, but that should only throttle the rate, not stop their scraping all together.
Clearly, places like Reddit have wised up to this and are making API usages non-free for example, so while it's not impossible, you can see the limitations being put into place already. Twitter is another one.
It seems like all this data is now considered gold and people lock up gold?
Why do you think chatGPT lost its Web Search plugin lately? Copyright lawsuits. You can't even use copyrighted content in the prompt because it will make the model makers liable.
Big publishers are more than happy to settle with AI companies - they just their slice of the pie after all. But who is going to protect, say, your Hacker News comments? Are you going to sue the AI company? Is YC going to sue? Are you going to sue YC for not banning crawlers in their robots.txt?
Heck, collecting whatever cents I might be owed for being a drop in the ocean is a losing move in my country.
Scraping & siphoning content, ad blockers, both sides have been in arms race for years, AI and LLMs will just be the last straw, before almost anything of value is behind a pay/subscription wall.
GPT is already capable of incredible generalized language understanding and I'd wager we've long since hit diminishing returns from raw internet data. RLHF, fine tuning, and better (and more data efficient) architectures are what we need now.
I'm sure you can do lots of interesting research using outdated datasets but for companies creating products this will not be sufficient.
If your data or models don't account for those then it can make mistakes. For example, if a model is only trained on modern sources (and does not know about Early Modern English 2nd person pronouns "thy"/"thine"/etc.) then it can easily get confused when determining parts of speech, which then affects other down-stream processing.
Would help to stay up to date on current events, and up to date on its general understanding of the world.
With the risk being adopting the bias of news media.
Nobody will make anything value in public, that they didn't want to be released for free anyway.
ChatGPT's vacuum has brought back a desire for privacy and will probably contribute to destroying piracy too.
ChatGPT has destroyed the 'study hard and get reward loop' for collaborative effort on the internet. If you use chatGPT, it absorbs all your question data and gives you nothing in return. You can't commit to random people, as they are expected to leak your IP onto gpt.
Isaac Newton using chatGPT would upload the core of calculus to GPT in research questions, and see no personal benefit for doing so.
There is no greater thief in history of academic work, than electronics.
across the political spectrum: every single one I checked, other than ft.com blocks GPTBot
For instance blogs will stop posting if chatbots in search grab all the answers directly from their sites bypassing all ad revenue.
I don't know why you would have that concern. LOTS of people write because they just want to share. Now broaden it out to include speech. LOTS of people talk - it's what we do.
Progress in AI will happen because humans like to express themselves. The challenge isn't copyright. It's figuring out how to capture the vast content that just isn't getting captured. Also, this is really only an issue for "new intelligence" - if you really think there is such a thing. Personally, I do not. I think like 99% of all human intelligence is in the out of copyright corpus.
The blogging scene from the early millennium is now a shadow of its former self, and one of the most often stated reasons for abandoning blogging is “my site just wasn’t getting many views any more”. In a world where AI generated content abounds, there will be even fewer eyeballs on whatever one shares and therefore less feeling of reward for sharing. Moreover, the people still blogging are often loading their content with referral links, because in an economy full of glamorous influencers, even ordinary people are tempted to seek some financial reward for sharing beyond the mere pleasure of it. Less eyeballs due to AI competition means fewer people clicking those referral links.
I'm not truly sure that llms mean less eyeballs, though. They produce mediocre content in an arena where high quality content matters. There's already a massive pile of crap on the internet; it's already all about surfacing the relevant and the interesting bits.
You just gave up the "for a living" group, who arguably produce overall better content (of course there are exceptions), and focused on hobbyists. I'd call that a self-defeat.
Humans generate a massive amount of natural language, and 99% of it never gets someplace where GPT/LLM training can consume it. If we can capture just a couple percent of that, then there will be no need for GPT/LLM to make use of content from those who don't want their writing to be consumed.
39,516,981,435 books3.tar.gz -- 36.5% compression ratio
108,371,325,720 tar -- uncompressed
Recompressed with 7-zip and xz:
25,221,357,605 b3.7z # with flags: -m0=ppmd (23.3%)
27,077,329,052 books3.tar.xz # with flags: -e9 (25.0%)
To see what other slower compressors could do, I checked results from a random 1,000 books. MCM would achieve about half the original tar.gz file size, but it's very slow.
104,043,048 q-x11.mcm - 19.0% -- from q.tar
106,181,190 q-m9.mcm
114,564,722 q.tar.bsc-m03-b1000000000
118,636,713 q.tar.bsc-m03
130,621,411 q.ppmd -- 23.9%
137,874,349 q-s29.tar.lzip
138,274,566 qultra.7z
141,815,136 q.lzma
142,626,724 q9.xz -- 26.1%
146,786,015 q.brotli
146,933,603 q.p7zip
146,933,619 q.7z
149,371,106 q9.bz2
198,213,051 q.gz -- gzip -9 -- 36.3%
198,213,193 q.zip
227,074,912 q.lz4
228,463,000 q.lzop
546,201,600 q.tar
ROT13 the text is the data still a copyright violation? I imagine so since it's a trivial thing to restore it to a legally volatile form.
I understand that an unresolved issue is whether, once ingested into an LLM, the trained LLM is in violation of copyright. One wonders if human readers too are in violation of copyright for having been "trained" as well when they read a book.
Is a "brain transplant" from one LLM to another a thing? Perhaps just a trivial copy of the node weights or whatever they're called. That would would not let the target LLM off the hook with regard to copyright violation I expect.
But what if one LLM "taught" another. Maybe that is not a thing yet.
At least it's a thing for vision models and was also used to train OpenAI Five (the dota AI)[0]. So it probably applies to LLMs too.
[0] https://cdn.openai.com/dota-2.pdf (page 2)
I also suspect that if we had an effective mechanism to prevent use of copyrighted work in training, it would necessarily behoove an artist to opt their content out. Will you really want to be excluded what may well become the canonical mechanism for searching and generating language?
LG literally has this info on its front page. Researchers will research, I suppose.
Still, there’s always room for books4.
These models are going to come out so fast from one-off contracts. Its not the line that some creatives think it is. Its a 2 year delay at best.
It isn't as though the AI companies have even paid for a single copy of the authors' books.