OpenAI and journalism
openai.com
openai.com
ChatGPT can replace the NYT archives (a paid service?) though.
Using public facing data for training is exactly as legal as looking at that data yourself. An AI regurgitating information is exactly as legal as a human doing it. If you are allowed to - in a public square - recite from memory an article you read on the internet... if that is legal: then an AI doing the equivalent must also legal.
This is a truth: because (to reiterate a basic fact) the training data (the text) simply does not exist in the LLM. What does exist is a set of weights and biases that are a representation of that data (a model). To say a model even is reciting data is inaccurate...
With that said - big companies are not your friend - open source development of AI is the only way to go. Transparency is key to healthy development.
You are in fact not allowed to do that
> You are in fact not allowed to do that
You're arguing poetry “recitals” aren't legal?
Copyright law is designed to protect original work from being used in a way that could harm the creator's interests, typically economically. In the case of a public recital, we have:
1. Fair Use presuming it's a one-time event, non-commercial, with no impact on market value
2. Public Performance generally means plays, music, movies, things that are a "performance". Reciting the article isn't the performance of something designed to be a performance, in fact it's arguably transformative.
3. Copyright infringement usually concerns reproduction, distribution, or commercial exploitation. A random, one-time, live recitation doesn't easily fall into these categories, and enforcing copyright in such a scenario would be impractical and unusual.
Certainly there are ways this could go differently. If the recital were to be recorded and distributed, or if it were done 9 to 5 daily as a "news barker" reading every top article from a soapbox as an alternative to the newsstand nearby ... could be a problem.
No.
I think the strongest point you make is the one-time thing. But I still think, in general, simply reproducing an article via speech verbatim would not be found fair use (not saying whether it should or shouldn't be, nor whether this at all applies to GPT, nor whether it's right or wrong, nor whether practically anyone would face consequence for it)
The key point is - this isn't being copied from the website, it's being generated from the model. this is not copied copywrite work being copied. What that is is speech. Is it protected speech?
Yes, it is.
It was copied from the website into the models training set, compressed into the model, and then reproduced from the compressed form.
> What that is is speech.
Even if it was, the Constitutional protection on speech vis-a-vis the copyright power is coextensive with statutory fair use (a codification of Constitutional case law), so that assertion doesn't, even if accepted as meaningfully true, change how copyright law applies.
The website is not being copied to an internal model anywhere, it is being looked (processed through an LLM) at and trained on...
If you can find it on the (open and public facing) internet, then it's fair game to everyone (and everything) to 'see'.
The quotes around “seen” are appropriate since its at best a metaphor.
A literal copy is made of the data (along with other data) to create the training set, and then a process is run which creates a lossy compressed representation of all the data in that training set, and that lossy compressed representation is called a model.
> there is NO text from any website in any LLM model.
There is, just as if you make a collage out of a series of pictures and then apply a lossy image compression algorithm visual data from those pictures is included in the compressed reperesentation.
> The website is not being copied to an internal model
This is simply false. That's exactly what is happening.
> it is being looked (processed through an LLM) at and trained on...
That is (lossily) copying into the trained model.
> If you can find it on the (open and public facing) internet, then it's fair game to everyone (and everything) to 'see'.
It is not free to those who see it to reproduce it; public display of a copyright-protected work doesn't waive protection, and even where seeing is a literal rather than metaphorical description, as for a human reader of a protected work, reproduction of the work or creationnof derivative after having seen it is prohibited.
The model is a (lossily) compressed form of the training data.
Following your line of reasoning if it is perfectly legal to walk into a coffee shop and sit down and listen to what the people next to me are talking about, commit it to memory, even make notes about it, does it then follow that it should be perfectly legal, reasonable, and acceptable for a govt agency or some other organization to put microphones everywhere to record what everyone is talking about, then feed all this data into various databases and modeling systems?
Reciting something in a park is different than selling a copyrighted print of something in a park when you don’t hold the copyright. Which is much closer to what the NYT is accusing OpenAI of.
The training data not “existing” in the model is interesting, but at some point, a distinction without a difference.
If I hire an autistic savant to go to a library and read all the books, then I set up a book selling service where whenever people want to buy a book I have my savant employee type out the book for them, is it then going to pass muster in a copyright case if I tell the judge “It’s okay actually, because the books don’t actually exist in my employee’s brain, merely neuronal encodings of them.” ?
If I have a copyrighted image on which I don’t hold the copyright. But I want to start selling it to people, is it cool if I just run it through a lossless compression algorithm, thereby generating a new encoding of the information and then sell this new encoding along with the software and command to reverse the compression?
Regarding the open source stuff, there I think you might find more favor to your arguments.
But the stuff we are seeing within commercial enterprises like OpenAI and Midjourney is clearly copyright infringement.
And I don’t see copyright law being insane in these cases.
As far as the savant reading all the books analogy goes... it's a bit off base - mostly because the AI isn't attempting to do that... it would have to be prompted specially (which as far as I understand - what's happening: people giving verbose special prompts to 'extract' copyright... which again... extract - the verbiage regenerate is better, considering there's no guarantee the generation will be a perfect reproduction...) to generated that information. What is happening, (fixing the analogy) the savant reads all the books in the library: then someone asks him to generate a brand new book... which contains some passages that happen to be like those in copy-written works... this is 100% interoperable to what human writers do all the time. Why would we ever want to punish an AI for reading and remembering better than us?
On top of that imperfect reproduction is the sale as if it's the original... that's a lot of additional assumptions to make...
Sadly the lossless compression is also a bad analogy. Math maps and 100% translatable and thus not change/encoding to the bits... if you compresses it lossy, to the point of doing it artistically, then... if none of the bits are the same - it's not the same picture, and doesn't hold any 'bit' of the old image.
Good reply!
No, and I do think OpenAI returning copyrighted works verbatim is probably copyright infringement even if it’s “laundered” through a LLM.
However if the autistic savant only provided summaries, analyses, etc that is fair use (IANAL), and should be for LLMs too.
That probably means LLMs will need some sort of scrubbing process to ensure exact training data can’t be reproduced, or if that’s not feasible then some type of output filter that looks for training data (although that would be a problem for open source models)
This is a distinction without a difference if the text can be reliably reproduced. For example, a trivial neural network that always outputs the text is just a fancy encoding of that text.
I don't know enough about amnesia to know if humans even separate these things or not.
If they are separable, LLMs can be made much smaller without loss of comprehension, as "mere facts" can easily be punted off the network and into a pool of searchable documents.
Without that capability, we'd have to teach the models that verbatim quotes are not allowed, just as we have to tell each other about copyright law because humans can recall things sometimes — and in mere conversation it's considered unremarkable (or even good!) if our mouths or hands reproduce perfect reproductions. Poems, pledges, anthems, quotes…
…but for anyone who has memorised, say, Star Wars, you're forbidden from rather than incapable of writing down the entire script.
Well, if that wouldn't get you out, we wouldn't have Youtube, Google, Dropbox and any other site that probably could output copyrighted material. This would be the end of the internet idea.
It would actually be more in line with the “Internet idea” of a decentralized network, not a massive hub and spoke arrangement.
Saying, “Hey, we already have some big corporations where copyright infringement plays some role in their business model, why not add a few more in the form of “Open”AI, and whoever else.” Is not a good argument.
The centralized server farms and behemoth corporations are in many ways representative of what the commercial internet has become, but are not a fulfillment of the original “internet idea.”
If Google, YouTube, Dropbox didn’t exist you could still share files with your friends.
Sharing a file with a friend and him benefiting in some way from that is a different thing entirely than massive centralizing corporations flouting copyright rules to benefit commercially and increase their power.
More of the former, less of the latter.
OpenAI is saying they are going to try stop people tricking them into a copyright violation.
Should it be illegal to gather certain bits of meta data on NYT articles like:
* word counts
* word frequency
* sentiment analysis
* grammar and spelling
* facts about the worldLLMs don't work that way.
* They allow an opt out for training – but it was introduced in August 2023 and they don't retroactively clean up their model once you do opt out.
* Regurgitation is a rare bug – doesn't mean it isn't a problem, or one they cannot be sued over.
* Prompts that get ChatGPT to regurgitate are "not typical or allowed user activity" – this is the weakest excuse of all. It is their responsibility to control user behavior on a system they built. If I can type something in and get the response I want then I am using it in the intended way. Otherwise nothing on the internet can ever be illegal, because hey the user was the one who typed the URL.
And given how buggy OpenAI is in general, can’t trust them to actually honor this opt-out, nor protect it from abuse (e.g. they scrape copied content). They’ll need a model like YouTube’s scrubber that either deletes content or monetizes it on behalf of the copyright owner. And since OpenAI is making lots of money here, yes there is a technically feasible path towards royalties. But sama thinks he’s cool af and e/sigma blah blah blah.
https://community.openai.com/t/chatgpt-occasionally-reuses-t...
It can be illegal to access material that was legal to host, because a user and a host can be in different jurisdictions.
It can also be possible to use a tool in an unlawful manner even if user and supplier are in the same jurisdiction, for example you can type up a document that reproduces a work protected by copyright without permission and print it out, and this is on you not on whoever you got the keyboard, computer, and printer from.
These may well not be sufficient defence in law, I am not a lawyer and can't tell what the law considers sensible.
Cast #2 – I host a blank database. You fill it up with copyrighted data, and later access that data.
You are alluding to case 2, and yes over there it is debatable who committed the violation. You will mostly likely be at fault if it fits under section 230.
This situation however is case 1. OpenAI is the one who accessed the copyrighted data, copied it, hosts it, and is giving everyone access to it. There is zero argument that the user is at fault here (which is what they are trying to say in this blog post).
It's supposed to be, but it regurgitated the NYT article and that's the problem.
ChatGPT is the one producing the article and profiting off it.
If ingesting data is a fair use violation it doesn't matter what the model outputs. I was responding that is the responsibility of the program creator stop all fair use violations instigated by its users.
2. Copyright. If it's a right, it's a right.
3. This is not a universally agreed upon position.
Of course, they don't want to do that.
You do realize that GPT-4 was a moment similar to the first flight of the Wright brothers? You don't have Boeing standards of quality yet.
I get the sense that there would be more backlash against these models which would drive us more quickly towards a lasting resolution if people better understood how their data is being used.
I don't think that training on copywritten data is necessarily wrong, just pointing out that doing so within the context window rather than at weight training time might not be so different.
It could potentially do similar searching by the internet for similar things to your question and then figuring out how to derive the answer, without finding an exact answer match directly.
I'm not sure who this post is for.
The normals who just use the ChatGPT webapps and mobile apps don't care about the lawsuit anyways. In that case, this blog post is just a Streisand Effect.
If changing your processes to include OpenAI is expensive then that’s a big concern.
Wouldn’t surprise me if their sales people are getting questioned on if the lawsuit is an existential threat and this is an attempt at giving them something.
Their legal defense will come in the form of a response to the complaint and a motion to dismiss.
So what does this mean for all content trained before August 2023?
Are you only allowed to opt out for future training?
Their data collection is very much automated, so avoiding opted out content that was reposted somewhere else is likely impossible for them.
Training on anything should be an explicit opt-in process and they (and every other model out there) should be retrained from the ground up with only explicitly allowed materials.
> Google's unauthorized digitizing of copyright-protected works, creation of a search functionality, and display of snippets from those works are non-infringing fair uses. The purpose of the copying is highly transformative, the public display of text is limited, and the revelations do not provide a significant market substitute for the protected aspects of the originals. Google's commercial nature and profit motivation do not justify denial of fair use.
IMO this is very different from what OpenAI is doing.
- The purpose of Google books was to search for content and then use the results of the search to buy the books from the original publishers. OpenAI doesn’t attribute results nor provides a way to pay for the content it returns.
- Google books doesn’t provide the entire book as a response from the query, only the snippet with the result. OpenAI does return entire articles verbatim or paraphrased as result of queries.
- Users of Google Books do not use it in place of actual books. ChatGPT replaces the need to read the actual articles/sources.
I wonder how they operationalize "meaningful". After all, if the work didn't meaningfully contribute to the training of existing or future models, then why do they seem to need it in the first place?
Relative to the rest of the data, NYT's is not meaningfully impactful. But they need the aggregate amount of data, which is why they need training to be fair use and to be opt out rather than opt in.
So arguably the importance of including the NYT and everyone else collectively in opt out selection is precisely because it makes the NYT individually less important in the final model.
What if I paste in some of my own code for ChatGPT to review and it's so terrible it barfs? Does that mean I violated the terms of use?
https://donhopkins.medium.com/the-future-of-gpt4-programming...
Fair use only applies to works you can use; it doesn't force their original authors/owners to publish them in the first place.
I don't see how that couldn't appease copywrite holders, shy of them actually just wanting to cripple threats to their old ways industry.
Proprietary software does not empower people.
- The purpose and character of the use
- Whether such use is of a commercial nature
- The amount of reference material used
- The effect on the work's value
I'm very much pro-AI but I'm not willing to sacrifice the fourth estate or erase the value of intellectual labor so some guy can by another half-million dollar watch.
I hope the NYTimes succeeds in their lawsuit, not because I want to hurt OAI or because I like the Times, but because I think the more we commoditize information, the dumber we get.
What’s the use case here NYT is trying to prevent? A user spending hours to force the model to regurgitate articles that may or may not be hallucination? Sorry to say it but there are far easier ways to get around the NYT paywall.
Pacifying the NYT isn’t worth it if it means we lose models like GPT4. I hope OpenAI wins this one.
NYT isn’t the only source affected by this. All content created and posted on the internet is subject to being used by ChatGPT. Their goal is to provide answers to any question you might have, as an alternative to browsing actual sources.
If OpenAI wins and ChatGPT continues to improve, there will be no financial or social motivation to create honest content for the internet. The death of the organic internet will accelerate and all content that remains will be created by bots.
Of course this is a dystopian view that may not fully realize, but if all it takes to avoid it is preventing a single corporation from succeeding, then I’m all for it.
You mean like the internet browsing became an alternative to actually buying newspapers? And now LLM answers an alternative to actually browsing?
Let's close both the internet and AI and save journalism. Because it's gone to shit.
It's a transformative application of copyrighted materials it's the textbook definition of fair use, and anyone pretending it isn't because of some anti generative AI luddite sentiment is risking to destroy a precedent on which much of the modern internet is built.
Question: Does this apply to paywalled content too ?
All GPT4 is doing is writing about the act of your writing, same way a journalist only creates from the act of another, would make sense if the individuals the journo writes about got paid a percentage for the act that caused the writing.
Seems hypocritical to claim the chain of writing about ends with you when you have nothing without others first acting when the page was blank.