In the case of The Pile, "processing for language modelling" means "converting epub and pdf into plain text, maybe deduplicating, maybe removing some sorts of detectably malformed files"
So not a particularly lossy conversion.
Then "they're doing something amazing, they don't need permission, and the cat is already out of the bag, and similar musings".
Seriously, it's both copyright infringement, and unethical. This is why I don't use any of the popular AI tools, or even AI add-ons in Evernote, Notion, etc. They all link back to the usual suspects.
Ah, the Uber theory of law. Works surprisingly well for some reason.
Open science repositories would take down the "dataset" immediately (or at least limit its access) if a copyright holder brings the matter to the eyes of the admins.
I can see the problem where direct and faithful replication is possible but where it isn't is there still a problem? Or is the automatable aspect, the scale at which it can occur, that is the problem?
An AI system consumes something perfectly, then ingrains it into its weights perfectly, and becomes capable of imitating the same thing perfectly. Plus, ther are no other internal or external factors which affect these "generation" over time. Hence, it mixes and reproduces based on what it consumed, solely.
I might get inspired by people, and add my own values to it, iterate over it and diverge from what I'm inspired from ultimately to create my own style. AI doesn't work like that. Also, if I do the same amount of inspiration with the same precision and accuracy, I'll be neck deep in accusations and lawsuits (for the right reasons).
As a result, just because we fail to ask the right questions to reproduce the training data verbatim or almost verbatim doesn't mean that the information is not there. At the end, a neural network is a compression algorithm which encodes data in terms of weights. Given the correct input, you can regenerate the training data as is.
Unless you have special abilities, you can't read 500 books an hour, remember them perfectly, and generate derivative works by mashing all of them together. If I do and try to sell a novel, I'll be ridiculed no end. If I write a Ph.D. the same way and try to defend it, I'll be banned from academia for three lifetimes at least.
For more elaboration on the subject, see [0].
If you divide Stable Diffusion's file size by the number of images used to train it, you get something like 1.2 bits per image, and it is physically impossible to get this kind of a compression ratio.
The actual problem with AI is that it sometimes plagiarizes random fragments of the work it is trained on, even if it is not the user's intend, and we currently don't really know how to fully prevent this.
Same for code generating models trained on Open Source and Free Software. Tons of licenses violated, from strong copyleft to source available models and reproduced (almost) verbatim with comments intact.
Some researcher's codebase is almost completely reproducible without any licensing information just by hinting the function names.
Maybe for the image compression it's borderline impossible for now due to network size, but for text and code, generation of training data almost verbatim is very possible and straightforward.
Also in image generation models, style transfer is the bigger problem, because it completely eliminates the artist who uses/created the style in the first place. "You pioneered this, and we fine tuned this model with your images, and we can do your work for free, without you, have a nice day". However, the artist's life expenses doesn't disappear when they're transferred to an image generation model.
This is also unethical.
Just because you can doesn’t mean you should. That's what I'm trying to say.
IIRC it was like 1.4 bytes before adding in random initial and prompt. And Amiga Four-byte Burger is 4 bytes long.
[0]: https://bytecellar.com/2023/04/24/lost-amiga-four-byte-burge...
In terms of fair use, one of the larger factors is the 'market substitution' factor, which basically means "does this use compete with otherwise licensed uses that people would ordinarily pay for?" AI absolutely does compete with human artists for the same market. In fact, it's winning handily[1], because you don't have to pay human artists. AI art models absolutely shouldn't be trained on anything with copyright on it.
The other factors don't fare much better. Nature of the original work will differ based on the plaintiff, but the purpose and character of the AI's use of that work is very much commercial. And the amount and substantiality of the use is complete and total. I don't see AI being fair use - at least, not in every one of the many, many training lawsuits currently ongoing against OpenAI and Stability.
[0] Starting with any body of text, an LLM, and an empty context window, compute the next-token probabilities and take the highest one. If it matches the source text, output a 1 bit. If it doesn't, output 0 followed by the ID of the correct next token. Add the correct token to the context window and repeat until the text has been fully compressed. This produces a list of perplexities (wrong words) for the given text which can be used to guide the LLM to output the original work.
[1] Hey, remember when both WotC (biggest art commissioner on the planet) and Wacom (hardware vendor that sells art tools and payment terminals[2]) both got caught using AI art after making very loud and public pledges to not do that? They both wound up buying stock photography on marketplaces that are absolutely flooded with AI trash.
[2] All the credit card readers in Japan are built by Wacom, which is really funny as an artist
And that's legitimate inventions I'm talking about. Just wait until the patent trolls figure out how to get the A.I.'s to divulge patent violations to sue the A.I. suppliers.
So in that sense I understand the response that "they don't violate copyright" by studying the material. Again, I don't pretend to be a lawyer, and not every law has to follow my logic.
This isn't about the output for content generators or about the abstract numeric weights that they operate over. That's more complex and a largely open question.
But this is literally about indiscriminately distributing copyrighted works in a large, convenient archive while arguing that it's okay because you normalized the formatting a bit and because you suspect that some people might find "fair use" value in it.
Meanwhile the complainant (such as GRRM) forgets that often passages from said articles and books are strewn throughout the Internet. Of course chat GPT can drop passages from the GoT books; there's several entire fucking wikis for that franchise that reference passages, quotes, details etc.
Same goes for news articles, passages of which are often quoted by other sources or websites.
Not that chatGPT has reproduced many works in whole, but it's an interesting logic problem for fair use law: if I have copyrighted article X, but websites ABCDEF all quote various passages of my article (ie fair use, critique, etc) and then ABCDEF is used to train an LLM, if the LLM can _reassemble_ the article from quoted passages without referencing article X itself, is it copyright infringement or fair use?
From the ruling:
> Assuming the truth of Plaintiffs’ allegations - that Defendants used Plaintiffs’ copyrighted works to train their language models for commercial profit - the Court concludes that Defendants’ conduct may constitute an unfair practice.6 Therefore, this portion of the UCL claim may proceed.
https://caselaw.findlaw.com/court/us-dis-crt-n-d-cal/1158180...