ChatGPT can replace the NYT archives (a paid service?) though.
Using public facing data for training is exactly as legal as looking at that data yourself. An AI regurgitating information is exactly as legal as a human doing it. If you are allowed to - in a public square - recite from memory an article you read on the internet... if that is legal: then an AI doing the equivalent must also legal.
This is a truth: because (to reiterate a basic fact) the training data (the text) simply does not exist in the LLM. What does exist is a set of weights and biases that are a representation of that data (a model). To say a model even is reciting data is inaccurate...
With that said - big companies are not your friend - open source development of AI is the only way to go. Transparency is key to healthy development.
You are in fact not allowed to do that
> You are in fact not allowed to do that
You're arguing poetry “recitals” aren't legal?
Copyright law is designed to protect original work from being used in a way that could harm the creator's interests, typically economically. In the case of a public recital, we have:
1. Fair Use presuming it's a one-time event, non-commercial, with no impact on market value
2. Public Performance generally means plays, music, movies, things that are a "performance". Reciting the article isn't the performance of something designed to be a performance, in fact it's arguably transformative.
3. Copyright infringement usually concerns reproduction, distribution, or commercial exploitation. A random, one-time, live recitation doesn't easily fall into these categories, and enforcing copyright in such a scenario would be impractical and unusual.
Certainly there are ways this could go differently. If the recital were to be recorded and distributed, or if it were done 9 to 5 daily as a "news barker" reading every top article from a soapbox as an alternative to the newsstand nearby ... could be a problem.
No.
I think the strongest point you make is the one-time thing. But I still think, in general, simply reproducing an article via speech verbatim would not be found fair use (not saying whether it should or shouldn't be, nor whether this at all applies to GPT, nor whether it's right or wrong, nor whether practically anyone would face consequence for it)
The key point is - this isn't being copied from the website, it's being generated from the model. this is not copied copywrite work being copied. What that is is speech. Is it protected speech?
Yes, it is.
It was copied from the website into the models training set, compressed into the model, and then reproduced from the compressed form.
> What that is is speech.
Even if it was, the Constitutional protection on speech vis-a-vis the copyright power is coextensive with statutory fair use (a codification of Constitutional case law), so that assertion doesn't, even if accepted as meaningfully true, change how copyright law applies.
The website is not being copied to an internal model anywhere, it is being looked (processed through an LLM) at and trained on...
If you can find it on the (open and public facing) internet, then it's fair game to everyone (and everything) to 'see'.
The quotes around “seen” are appropriate since its at best a metaphor.
A literal copy is made of the data (along with other data) to create the training set, and then a process is run which creates a lossy compressed representation of all the data in that training set, and that lossy compressed representation is called a model.
> there is NO text from any website in any LLM model.
There is, just as if you make a collage out of a series of pictures and then apply a lossy image compression algorithm visual data from those pictures is included in the compressed reperesentation.
> The website is not being copied to an internal model
This is simply false. That's exactly what is happening.
> it is being looked (processed through an LLM) at and trained on...
That is (lossily) copying into the trained model.
> If you can find it on the (open and public facing) internet, then it's fair game to everyone (and everything) to 'see'.
It is not free to those who see it to reproduce it; public display of a copyright-protected work doesn't waive protection, and even where seeing is a literal rather than metaphorical description, as for a human reader of a protected work, reproduction of the work or creationnof derivative after having seen it is prohibited.
The model is a (lossily) compressed form of the training data.
Following your line of reasoning if it is perfectly legal to walk into a coffee shop and sit down and listen to what the people next to me are talking about, commit it to memory, even make notes about it, does it then follow that it should be perfectly legal, reasonable, and acceptable for a govt agency or some other organization to put microphones everywhere to record what everyone is talking about, then feed all this data into various databases and modeling systems?
Reciting something in a park is different than selling a copyrighted print of something in a park when you don’t hold the copyright. Which is much closer to what the NYT is accusing OpenAI of.
The training data not “existing” in the model is interesting, but at some point, a distinction without a difference.
If I hire an autistic savant to go to a library and read all the books, then I set up a book selling service where whenever people want to buy a book I have my savant employee type out the book for them, is it then going to pass muster in a copyright case if I tell the judge “It’s okay actually, because the books don’t actually exist in my employee’s brain, merely neuronal encodings of them.” ?
If I have a copyrighted image on which I don’t hold the copyright. But I want to start selling it to people, is it cool if I just run it through a lossless compression algorithm, thereby generating a new encoding of the information and then sell this new encoding along with the software and command to reverse the compression?
Regarding the open source stuff, there I think you might find more favor to your arguments.
But the stuff we are seeing within commercial enterprises like OpenAI and Midjourney is clearly copyright infringement.
And I don’t see copyright law being insane in these cases.
As far as the savant reading all the books analogy goes... it's a bit off base - mostly because the AI isn't attempting to do that... it would have to be prompted specially (which as far as I understand - what's happening: people giving verbose special prompts to 'extract' copyright... which again... extract - the verbiage regenerate is better, considering there's no guarantee the generation will be a perfect reproduction...) to generated that information. What is happening, (fixing the analogy) the savant reads all the books in the library: then someone asks him to generate a brand new book... which contains some passages that happen to be like those in copy-written works... this is 100% interoperable to what human writers do all the time. Why would we ever want to punish an AI for reading and remembering better than us?
On top of that imperfect reproduction is the sale as if it's the original... that's a lot of additional assumptions to make...
Sadly the lossless compression is also a bad analogy. Math maps and 100% translatable and thus not change/encoding to the bits... if you compresses it lossy, to the point of doing it artistically, then... if none of the bits are the same - it's not the same picture, and doesn't hold any 'bit' of the old image.
Good reply!
No, and I do think OpenAI returning copyrighted works verbatim is probably copyright infringement even if it’s “laundered” through a LLM.
However if the autistic savant only provided summaries, analyses, etc that is fair use (IANAL), and should be for LLMs too.
That probably means LLMs will need some sort of scrubbing process to ensure exact training data can’t be reproduced, or if that’s not feasible then some type of output filter that looks for training data (although that would be a problem for open source models)
This is a distinction without a difference if the text can be reliably reproduced. For example, a trivial neural network that always outputs the text is just a fancy encoding of that text.
I don't know enough about amnesia to know if humans even separate these things or not.
If they are separable, LLMs can be made much smaller without loss of comprehension, as "mere facts" can easily be punted off the network and into a pool of searchable documents.
Without that capability, we'd have to teach the models that verbatim quotes are not allowed, just as we have to tell each other about copyright law because humans can recall things sometimes — and in mere conversation it's considered unremarkable (or even good!) if our mouths or hands reproduce perfect reproductions. Poems, pledges, anthems, quotes…
…but for anyone who has memorised, say, Star Wars, you're forbidden from rather than incapable of writing down the entire script.
Well, if that wouldn't get you out, we wouldn't have Youtube, Google, Dropbox and any other site that probably could output copyrighted material. This would be the end of the internet idea.
It would actually be more in line with the “Internet idea” of a decentralized network, not a massive hub and spoke arrangement.
Saying, “Hey, we already have some big corporations where copyright infringement plays some role in their business model, why not add a few more in the form of “Open”AI, and whoever else.” Is not a good argument.
The centralized server farms and behemoth corporations are in many ways representative of what the commercial internet has become, but are not a fulfillment of the original “internet idea.”
If Google, YouTube, Dropbox didn’t exist you could still share files with your friends.
Sharing a file with a friend and him benefiting in some way from that is a different thing entirely than massive centralizing corporations flouting copyright rules to benefit commercially and increase their power.
More of the former, less of the latter.
OpenAI is saying they are going to try stop people tricking them into a copyright violation.
Should it be illegal to gather certain bits of meta data on NYT articles like:
* word counts
* word frequency
* sentiment analysis
* grammar and spelling
* facts about the worldLLMs don't work that way.