New York Times considers legal action against OpenAI as copyright tensions swirl
npr.org
npr.org
It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement.
While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accepted at scale, copyright in practice will change completely. In effect, the choices are "block this, or entirely destroy copyright protections in many industries". You can't allow this without eventually allowing everybody to simulate their own NY Times reporters, produce their own Marvel movies, and create their own Taylor Swift albums.
If you do allow that, the many many affected industries have catastrophic problems.
Problematic though copyright laws are, I see no world where all those protections go away any time soon, and so if the courts don't agree to protect copyright already in this scenario, then it will eventually be legislated to make that happen. AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken.
There's a grace period now while we work our way there, but the politics is pretty clear and with no plausible path to "let's drop copyright completely" ASAP, I just don't see any other result in the medium term. Doesn't mean the end of generative AI by any means, just a slowdown as we move to a world where you need to negotiate rights and buy data to feed it first, instead of scraping everybody else's for free.
Do you think Bing or Google are going to negotiate copying rights with the world's websites?
LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents and trademarks haven't really helped as much as they were expected to. Only recently did works as far back as 1920 enter the public domain. Additionally, some LLMs gobble up source code with mixed licensing structures. How is that to be handled?
Patents are about monopolies to produce something to hit a market, with the trade-off of showing everyone how it's made.
Trademarks just allow you to defend your name(s).
Copyright prevents other people from making money off of your work. For the rest of your life, plus 75 years.
Something is broken here, alright. While we're discussing intellectual property, what about one's DNA? Is it not a performance of biology? How about your fingerprint? Fingerprints are semi-unique, so it's also a performance mark. We've seen celebrities sue for the use of their likeness, so that's recognized to some degree as well. At what point will data subjects get rights, so that when another Equifax happens, they can be bankrupted and prevented from harming the public again?
Most are rhetorical, of course, but I really think generative language models are disrupting a lot of things we used to take for granted, and our models of creatorship are not refined enough to account for digital or statistical copying. Intellectual property as a concept is not compatible with a digital future.
Search engines index the web and point you at other people's work, along the way showing perhaps too much of that content (thus "stealing" users from the target webpage). But they don't reshuffle existing content into something apparently new and original.
The "malicious" case for generative AI is that it sucks in copyrighted work (vs. indexing it), rehashes and produces something that is supposedly original, but really a sophisticated rehash of copyrighted work.
As it stands right now, yes, Fair Use. But where do we draw the line? That line's been blurry for a while. Mostly limited to no more than 30 seconds of a performance, and no more than what's needed to quote literary works. Not sure about lyrics.
The issue is on some level, most things we create are derivative. Someone had to have the idea first, but once an idea is unleashed upon the world, it seems very difficult and unwieldy to put the genie back in the bottle.
I'm not really pro-LLM since it enables business to leech off of FOSS even more efficiently. The disruption of LLMs seems to be accelerating our philosophical re-examining of copyright and licensing terms in FOSS. We will inevitably need a GPL in the future that disallows remixing via LLM or other generative text algos, due mostly because ensuring all of a result is freely usable is not easy. Limitations could be built in to only 'fetch' code licensed under permissive terms, like MIT or BSD, but the tendency of these models to 'hallucinate' means you really cannot know the legal standing of LLM-generated code, at present.
LLMs do.
It's hard for me to reconcile that it's somehow OK for a student to write a term paper about something (e.g. "Ulysses"), Wikipedia doing the same thing, but on the other hand, not OK for chatGPT doing it.
No, they won't. Search engines have already fought and won this battle on fair use grounds because they make use of the copyrighted content differently than LLMs do. It's an absolutely fundamental distinction.
Patents and trademarks haven't really helped as much as they were expected to.
Patents have been a thing for nearly a millenia; Britain's patent system is credited with giving it the technological edge over its Medieval and post-Medieval competitors for world domination. Similarly, the U.S. patent system has been credited with the U.S.' technological prowess for the past 2 centuries.
Only recently did works as far back as 1920 enter the public domain.
Nothing is stopping artists from releasing works into the public domain during their lifetimes. That would of course would mean other people economically exploiting their works however they wished without any input or control from the artist...which is generally why most artists haven't done that.
While we're discussing intellectual property, what about one's DNA? Is it not a performance of biology? How about your fingerprint? Fingerprints are semi-unique, so it's also a performance mark
This is either a bad-faith argument. DNA and fingerprints are tangible things created without any sort of intellectual input. Therefore, by definition not intellectual property.
We've seen celebrities sue for the use of their likeness, so that's recognized to some degree as well.
This is another bad-faith argument. Celebrity likenesses are intangible property but are not intellectual property and are not afforded the same protections.
I really think generative language models are disrupting a lot of things we used to take for granted
LLMs so far have disrupted student papers. And that's pretty much it.
I’m curious if you can set me straight with a citation, or whether you would contest mine.
But ChatGPT is actually providing an alternative that obviates the original articles themselves.
Personally, I like the flexibility of an LLM being able to describe a process at different skill levels. This is of tremendous educational value to the world.
Even the link is a copyrightable item -- artistic effort went into creating it
Small snippets are allowed by copyright law. They are not infringing.
> and Google will show you their cached page if you ask for it.
Really? I haven't seen that in several years. How do you get it these days?
I always assumed Google quit giving you that option exactly because of copyright issues.
> Even the link is a copyrightable item -- artistic effort went into creating it
IANAL, but I'm pretty sure that the current state of copyright law disagrees with you. Can you point to some concrete evidence that you're right?
http://webcache.googleusercontent.com/search?q=cache:www.hac...
etc
There's also a link in the three-dot menu next to the search result, but it doesn't always appear.
LLMs output a mashup of source material without attribution. This is 99% of the time against the copyright holder's interest.
It’s basically the “does this replace the original content” doctrine of fair use
What you're thinking of is "featured snippets". As far as I know, the justification behind those is that they are exact quotes that are followed by a citation (a link). Google argues those are fair use, since it's a properly referenced quote.
Copyright largely remains about the PRODUCTION of content, not about the CONSUMPTION of it.
Someone who grew up reading Marvel comics being able to make new original comics in that style is perfectly ok. That same person perfectly replicating an Avengers comic is going to land them in hot water.
The focus on infringement really needs to be on what LLMs produce, not their training.
There definitely needs to be something like a secondary pass added which checks output against a vectordb of the training set to avoid too close derivative IP outputs (and ideally checks for jailbreaking or inappropriate content at the same time).
Any production services would need to subscribe to a service like that to stave off litigation on infringement, much like how the oft repeated complaints regarding YouTube copyright infringement eventually dissipated as content tagging was added (and shifted to complaints over too broad application of it).
A generative AI model having read the NYT but producing new original news articles in the style of a newspaper is a very weird argument for infringement.
A human driven service or an automated one that takes current NYT articles and summarizes or reworks them, publishing itself and cutting them out of the ad revenue is more problematic (but also widespread already and generally considered protected).
Services which exactly duplicate their articles would be more clearly infringement, but there's no evidence that's even a fraction of what ChatGPT is doing.
Criminalizing training would set back whatever county did so significantly in global competition for a critical new economic (and defense) trend, and would ultimately only be a minor stop gap for copyright holders as you'd simply see a market for secondhand generated content from foreign models trained on copyrighted data but then producing content that was itself not copyrightable but could be used to train domestic models.
This has napster -> subscription spotify energy. But the only people happy about that are Spotify and people who found it distasteful to download music illegally. There just wasn’t a consumer-friendly option for a while, so the black market was the only market.
So. The enforcement mechanism is what… a scary DMCA letter?
(There will definitely be a stupid DCAIA in the next congress)
In an (unrealizable) regime where all copyright holders are compensated, that would include picopennies for the discussion we’ve had!
There's a headline out today about several ex-Google Brain engineers, including a co-author of “Attention Is All You Need”, setting up shop in Tokyo. [1] That's not a coincidence.
> Amid rising questions about the fairness and legality of using publicly available information to train AI models, Japan affirmed that machine learning engineers can use any data they find.
> What’s new: A Japanese official clarified that the country’s law lets AI developers train models on works that are protected by copyright.
> How it works: In testimony before Japan’s House of Representatives, cabinet minister Keiko Nagaoka explained that the law allows machine learning developers to use copyrighted works whether or not the trained model would be used commercially and regardless of its intended purpose. [2]
IANAL so don't know what the implications of this are when it comes to cross-border copyright enforcement. It's hard to imagine Japan rolling back this type of legal safe harbor. It's a boon for attracting AI startups from elsewhere and giving the local tech industry a competitive boost.
What's to stop other venues looking to grow their tech industry to do something similar? And if they do, would it create a race to the bottom type of dynamic at the expense of copyright holders?
[1] https://www.bloomberg.com/news/articles/2023-08-17/ex-google...
[2] https://www.deeplearning.ai/the-batch/japan-ai-data-laws-exp...
OpenAI etc. have huge amounts of money behind them, they very well have a fighting chance in court to defend their usage of scraping the internet.
These creative industries include all of software development, music, TV, movies, books, media, art, etc. You do technically solve the problem of copyright by shutting all those down, but I'm not sure it's a solution anybody will vote for.
If you can come up with a serious alternative though, which can sustain those creative industries without requiring copyright, now is probably the best moment in all of history to seize the day and make that happen. There's going to be a big shake-up regardless, it's the perfect chance for alternative models.
Bear in mind that dropping copyright entirely doesn't just hurt Disney and Sony Music though - with no copyright the GPL and all other open-source licenses are unenforceable, anybody can copy & sell anybody else's art or design without permission, Spotify doesn't have to pay musicians even $0.01 any more, etc etc etc. It's not an easy problem.
That is the problem. Technically, AI should be allowed to 'read' content, it isn't hidden, and it gets mixed with other content in a 'brain' like thing.
AI and Humans can both spit out a new product that is 'similar' and thus be sued on that similarity.
But it can also produce endless similar variations at low cost and fast.
It is the ease of creating new similar products.
You could just as well prompt the AI "make a Taylor Swift song, but different enough to avoid a lawsuit".
Think this is an entirely new problem that needs a new law beyond copywrite. Copywrite is not the correct law to us for fighting this. Copywrite law doesn't ban someone from reading the source altogether.It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do.
It’s mechanically copying and reproducing part of its input data set. Making a tool that regurgitates others’ copyrighted IP is going to been seen as aiding mass copyright violations. Exactly like how Napster got sued: they’re holding a bunch of material they shouldn’t be. The only new twist to this case is the data is encoded in a transformer’s weights. This should be correctly seen as the same as having encrypted copyrighted data using a lossy algorithm.
This is not correct. AI models are tools that humans use.
This is like saying "it was typed on a computer therefore it doesn't enjoy copyright protections"
So now you’ve divided the world into those who use the best tech, and those who are not allowed. And that is what openAI wants, they’re betting courts are going to rule that however tainted the source, that we can’t put lightning back into the bottle.
Compare:
"Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/steve-jobs-def...
Against:
"Whether to describe SJ as a tyrant is a matter of perspective...": https://chat.openai.com/share/28633f0c-007f-48b6-a615-1581c3...
The general way LLMs work do not preserve content in it's original form: the ideas they contain are extracted and clustered statistically - as a ELI5 refresher, an LLM reads 2 million NY Times articles and records that after the word "Steve" there are a lot of "Jobs" followed by a lot of "was a genius/tyrant", "founded Apple", etc. Then LLMs recreate the user question "Who was Steve Jobs?" using this complex net of token/word stats. Is that fair use? I think OpenAI lawyers will not even tap the fair use question, they will simply state that no copy happened, just a statistical collection of words from various sources.
And importantly: no LLM source is really prevalent, so the end result cannot be even be traced back to the source, especially if multiple, similar news sources are being fed to training. I have no idea how the Times is going to prove that its _theirs_ news.
Sounds to me like they are trying to claim copyright over facts rather than the specific expression. That’s just not how copyright works at the moment.
The framing of openAIs recent changes is telling too. OpenAI seems to have nudged their models to reject requests of querying sentence continuation for specific sources - which the press is now framing as “trying to hide the use of copyrighted data”
What we are seeing here is an unprecedented attempt at expanding copyright doctrine to facts, style and information rather than specific expressions - a land grab of latent space by rights holders salivating to own factual information
Yea, now I can't read the paper and talk about it to other people it seems.
The Right to Read was a prophecy I guess?
The particular problem here is this program isn't magic, it just requires a lot of electricity and hardware to train at the moment. If at some point in the future this hardware becomes cheap then now suddenly OSS LLMs would be under the same set of rules that we're applying to major technology companies.
But mark my words, the large copyright holding groups don't give any shits other than how much IP they can scrape up and demand money for, for the next few human lifetimes.
Less glibly: a non-profit oriented LLM is just in a little different place on the scale, but doesn't fundamentally change my takeaway. However in this situation it makes it particularly egregious.
Which is hard, best hope they have is trying to put the burden of proof on the nytimes to show you can make the model regurgitate their articles (with some nudging).
If they manage that then nytimes is going to have a lot of trouble showing the model actually breaches their copyright, because just the information contained in their articles is not enough to constitute a copyrightable work.
The weights are also executable code (in some sense). When you query an LLM you're running this program with a given input. Yeah when it runs it tells a whole lot of things (sometimes novel combinations, sometimes verbatim repetition of trained data) but the point here isn't whether the output of the LLM is copyrighted; it's the weights.
You can make arguments like a) what is ChatGPT but a different kind of search engine, or b) what is an LLM but a primitive human, or c) but but uhh we didn’t agree to these terms.
But I do not think those arguments will prevail.
So if that’s the argument it’s already been argued by LinkedIn and lost.
This is one of those things where copyright holders have gotten absurdly full of themselves though. Like what you’ve said is that copyright holders have the right to impose a contract of adhesion on data that they are broadcasting into the public without any idea with whom they are even forming a contract, and that’s a facially absurd and incredibly noxious idea if you follow it to the conclusions it implies.
Copyright is about securing to the public works of significance and encouraging their creation and the way it’s become a lifetime-plus-75-year guarantee of intellectual ownership of ideas is fundamentally noxious and goes against the intent and spirit of the idea. And if that’s where the copyright regime is headed then I’d rather see chatGPT kill off copyright entirely.
A similar sort of issue popped up in the 80s around colorization of films. https://www.latimes.com/archives/la-xpm-1987-06-20-ca-8405-s... https://chart.copyrightdata.com/Colorization.html
The answer may be 'maybe'? As from what I read they basically split the decision down to 'i know it when I see it' style of ruling. If the copyright is still in effect then NYT owns that portion of the output but not others parts. As the secondary effect would be owned by the generator company (in this case OpenAI) or the person who prompted for it. If that is the case NYT would have to prove what parts (nodes? bacreferences? weights?) they own?
[0]: https://www.forbes.com/sites/zacharysmith/2022/04/18/scrapin...
I see they went to the Supreme Court who kicked it back to the Ninth who then re-affirmed their position that HiQ Labs was not in violation of the CFAA.
https://www.copyright.gov/fair-use/
Now, there's this idea that "news" is just factual and therefore falls under "fair use". However, that's only part of what section 107 says.
Fair use very much is still conditional, as there are 4 factors to be considered: (a) Purpose and character of the use, including whether the use is of a commercial nature or is for nonprofit educational purposes (b) Nature of the copyrighted work (c) Amount and substantiality of the portion used in relation to the copyrighted work as a whole and (d) Effect of the use upon the potential market for or value of the copyrighted work
The big issue isn't companies training LLM's using unlicensed materials (e.g. copyright protected works); it's publishing the output to the wider world. That's where a liability is created.
Is the way LLM work relevant? I can make a shitty script that has as input Microsoft proprietary code and as output something identical in purpose but the text is completely different, I would rename names with synonyms, swap some things around etc.
I am not against AIs, my opinion is that if your AI uses GPL code the output should be GPL, if it uses public domain images the output should be public domain images.
I mean for code if AI is actual intelligent you should be able to train an AI with C with just a few books and not with the entire GitHub open source code (and notice MS did not trained copilot on the proprietary code they have access proving they are not confident that they are in the right).
If I use Inkscape is the output of my drawing subject to the same terms as Inkscape?
If I use a Photoshop filter is the output subject to Photoshop's EULA and/or the copyright of the photo I started with?
If you get my image from the internet then you resize it in Photoshop you can't claim you created some original art, you just used the resize/crop/color filter function.
What ChatGPT produces under normal use is not more similar to the NYT source than any other article on the same topic.
ChatGPT is doing what a smart student does when he copies the homework, he combines a few sources and changes some wording. Technically there is no creativity, it is interpolating it's inputs and there is some randomness thrown in.
We also know that ChatGPT put some filters to filter out copyrighted outputs after they were caught that the AI actually memorizes paragraphs of text word by word.
https://en.wikipedia.org/wiki/Copyright_protection_for_ficti...
I am not sure why it would suddenly become infringement because an LLM is composing that email for me.
Not-A-Lawyer NAL instead of the full IANAL.
I just had to say it.
I fear that LLMs are going to cause the internet to be a much worse and less open space.
Let data be free! If someone wants to use it to make money, well, it's open, just like open source. It's still not okay to take open source work and claim it as your own, which is what copyright should be limited to. Stealing a photo or plagiarizing an essay is intrinsically different than just having a copy read by something, be it human or an mechanical process such as training a LLM.
OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills.
And I don't think it is fixable by making the AI act more like a search engine.
One could easily argue that it's unlikely it'll get worse - if anything, AI could empower competition as now a group of 3 passionate, free writers can compete with agenda-driven, for profit corporations on a similar level. This could very well make the web more open and free as it makes the web more accessible.
Now? Yes. I grew up when it was just ugly, and by ugly I mean animated gif backgrounds with obvious seams on the tile boundary.
*old man shakes fist at The Cloud*
Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media.
When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)
In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business.
I also hold a dislike for the copyright lobby, but this matter is serious. The question of whether the training of LLM's with copyrighted data is a legitimate one, as OpenAI did not just use contributions from large media outlets like NYT but capitalized on small contributions from individual contributors.
A ruling in favor of copyright would force OpenAI to shut down - but given their impressive demo of the tech, I hope we'd see more open and accessible versions of these models emerge.
As impressive as ChatGPT is, I dislike having my access to information governed by some large corporate entity. I also dislike a company directly capitalizing on my contributions without my explicit consent.
Who cares? If they want data, they can pay for it.
But in this sense, shouldn't they be suing Google as well? Since Google as a search engine, also crawls the web and shows their articles in their search results, usually it may even use them for those quick answers features.
My 5 cents on this are that NY Times noticed OpenAI has deep pockets, they may have ground to sue and decides to try their look in order to get some quick easy money. Now, I don't know if what OpenAI is doing with ChatGPT does not fall under fair use.
If, when someone reads a newspaper, they are served a paragraph-long answer from an NYTimes reporter that refashions reporting from local sources, the need to interact with the local sources is greatly diminished.
Very few people bother doing this for much the same reason very few bother fighting any of the other terrible decisions made by corporations with legal departments whose annual cost exceeds their personal lifetime earnings.
We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts.
Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay, physical locality, chemical delays and interference from things like myelin sheaths, and other physical processes affect the synapses that are partially mirrored by computer simulations of neural networks. Unlike neural networks, human brains also don't work based on a single clock signal triggering input and output from all notes in instant steps.
Human memories are also not just "data in, weights out". They are heavily modified by things like mood, concentration, language(s) spoken, context, and emotional triggers. There's no way to feed a dictionary into a brain. Memory preservation consists of multiple stages, with differing memory types, involving various brain segments with dedicated functionality that can actually grow back due to neuroplasticity in some cases.
Efforts are being made to emulate living cells on computers, but LLMs aren't that. Inversely, efforts are also made to feed brain cells artificial signals and train them to play video games, which results in different behaviour compared to the systems we use for LLMs or other AI systems.
It doesn’t seem clear to me that it does.
Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in a more simple mind/model?
What if the output is exactly the same, or comparable enough, regardless of the degree of complexity of the mind/model?
This would be a separate argument, though, from the notion I responded to above that the difference in the processes of a human mind and LLM are the reason why "learning" from copyrighted material is a violation of copyright in one case and not the other.
In my view the biggest issue to raise is the effect of the use upon the potential market for or value of the copyrighted work (https://en.wikipedia.org/wiki/Fair_use). The most spectacular example of damage would be stackoverflow, thought stackoverflow content is not copyrighted. I think there is little doubt that LLM's drive attention from original sources. That might be deemed damaging, especially in the long run.
There's no damage to the potential market or value of the works because they're given away for free by their owners.
There may be, because technically copy/pasting SO code is governed by CC BY-SA 4.0 license which requires attribution and things aren't so obvious especially for commercial purpose.
I’m sure there’ll be a human on the books, on paper.
If you turn an LLM into a person, you may have a ethical and legal basis for treating that LLM like a person. There's no law about artificial intelligence being sentient or not, but law applies only to people, so that'd be the supreme court case of the century. I remember the Star Trek TNG episode about this topic and while the answer was perhaps more obvious with mister Data, the best arguments for and against synthetic consciousnesses have all been made in that episode.
The law doesn't care for how human-like programmers may think their program is, and neither should it in my opinion. What matters to the law is that the output is a result of an automated process, which comes with a completely separate set of rules and conditions compared to fair use.
The complexity of the program isn't a very good legal defence in my opinion because there's no clear line when the complexity would be enough to be considered human like. You could, for example, also claim that a computer is just very good at doing imitations, just like a person can be good at doing imitations on stage, and that an mp3 file is just an elaborate imitation act.
Even with a digital system identical to a physical system I don't think you can state human-ness as an argument. A tape recorder is just a sophisticated, automated way of sending an electric field through a magnet, similar to what a human can do with a dynamo and a spool of tape; a sort of delayed-action theremin, which would turn it into a musical signal. In turn, a neural network can be solved by human brain power if you pay enough people to work on a single iteration for an entire year. Almost everything a computer can do is just a sophisticated way of doing what humans are already doing, so I don't see why this is different when it comes to this topic.
There's no obvious "this is human like now" threshold and I doubt there will be until we know exactly how the human brain works.
When it comes to AI generated works, we don't currently know where the line between copyright violation and copyright exemption lies. If this goes through, it's the third major lawsuit of its type, the other two being actions against Stable Diffusion by artists to prevent it from producing derived works from their art.
IANA but I think you can assume that the "but computers are just like digital humans" approach won't fly in court. That's not really important, though; both sides of the coin are already having clever lawyers write up legal defences for their points of view, and it's more than likely that some other factors ("is a model a derived work" (probably) and "is the output of a model a derived work" (who knows!)) will decide the future of AI and copyright. The interesting thing is that academic research is essentially exempt from copyright law, so nobody can demand takedown of their content from research data sets, but whether the commercial branch of AI companies can use their academically generated models to serve their customers?
As an upside, I don't think the death of ChatGPT is the end of AI. This whole scenario could've easily been avoided if AI companies paid for their data set or had restricted themselves to works they had the license to (public domain, CC0, etc.), and OpenAI in particular has been pretty brazen in their "we'll see about it if it ever comes up" approach. Companies like Github are probably in the best place legally, where their users are already signing off on "we can take your content and do whatever the fuck we want" terms and conditions.
I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to make even more money off of. We don't exist as computational nodes for them, a free resource to exhaust. It is a completely one way street with no consent. So I am in favor of all of these companies getting a reality check.
I think historically it's about copying wholesale and redistributing for profit. That doesn't seem to be what's happening here.
If I read a publicly available article, and I create a summary, is that covered by copyright? If I get an AI to do that, is that somehow a 'special' type of summary that is covered?
Can a provider of content somehow say "you may not use this for summarization"? Or apply other terms to my consumption if they are making it publicly available?
I think the comparison is that if you ask a person something like:
"Describe a power station?"
They will likely have never been in a power station. They will be leaning heavily pieces of content they have consumed over the years, that presumably were copyrighted. If you create an article that describes a power station, have you breached copyright and should be sued by all the people over the years that have produced content about power stations?
Is it suddenly different when a computer does it?
So you'd be fine with it if the model only ingested a humanly plausible amount of data? I suspect that would only make their legal issues worse, since the LLM would be much more likely to repeat tokens from the training set verbatim.
I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use.
Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't understand, but that the LLM essentially CAN understand.
Still a copyright abolitionist though. Maybe now more people will join the fight?
The weights are not a reproduction of the content. They are capable of it but so is a photocopier a lot more and we didn’t ban those either despite them technically being a lot more useful for violation.
Nah, this is expansionist doctrine and agenda for copyright - these companies are trying to copyright style and locations in latent space now.
It would be different if I trained my own AI, for my personal use.
Again, LLMs don’t copy so it’s not a good metaphor.
You brought up photocopiers and you never said it wasn't a good metaphor so I don't know why you pre-pended "Again".
If it's not a good metaphor, that invalidates your point that:
"we didn’t ban [photocopiers] either despite them technically being a lot more useful for violation"
So now you're arguing against yourself.
I'm going to bypass this question a bit and say, who cares?
Why do we need to treat these things the same way we treat humans? Why can we not say that it's okay if a human does it, and not okay if it's a computer? There's nothing that requires us to establish 'fair' as treating them the same as people.
We've been living in a world where people read news articles, then wrote almost the same article on their own website to sell ads. It's been a standard business practice for a decade now, what's so special about LLM based blogspam? The end impact is still the same, people reading the blogspam instead of the source
Because it would be absurd for it to be legal to do X, but illegal to do so with an efficient tool. Especially when the activity X in question is "learning".
It's legal in most places in the U.S. to own a semi-automatic weapon that fires x bullets/minute, but not a fully-automatic one that fires 10*x bullets/minute. It's legal for me to send water into my sewer by flushing the toilet, but not for me to pump the larger volume of water in my sump pump into that same sewer.
It's legal for a child to throw a handful of sand from the beach into the ocean. Do you think it should be legal for anyone to build and operate a sand throwing machine that methodically throws all of the sand on the beach into the ocean?
No such thing, "learning" in a vacuum describes nothing here. Might as well ask why I am allowed to make noise, e.g. speak, but when I install 5000 watt speakers on every square meter of the planet suddenly it's a problem, and roll my eyes at the inconsistency of not being allowed to "do X more efficiently with a tool".
Did we somehow stealthily develop a neural interface that lets us feed the 'learning' that 'AI' is doing into a human brain? Have we actually figured out how to do that?
No, we haven't. So humans are still learning the same way, but with a new tool to condense and summarize some information. Kinda like a textbook in school. But we don't treat those as human beings with human rights do we?
This isn't as well-defined for humans as you might think, so can't be well-defined for LLM techniques by comparing them to human agents.
My argument against the current hoovering up of data under various licences for AI training, which they claim can#'t reproduce anything verbatim, is CoPilot. If there is no risk, then why did they only use public repositories and none of their own private ones? Surely they think their code contains good training material, unless they think their own code is gobbledygook. Or back in terms of licences: if it can't breach the GPL family, then it can't breach their own commercial licensing arrangements.
Would OpenAI have a problem with humans for using some of their code/documents/other in this way?
So I can write my own article about the sky being shown to appear blue much of the time, but I can't copy someone else's article about the same subject.
Irrespective of copyright issues, the question is how to avoid creating a new class of large rent seekers in the LLM space.
In the short term, of course, the existing law matters, but the main discussion should be not on how to apply existing law but how to ensure that the new laws match what we-the-people would want.
I think when this happens it is normally easier to block a law than to push it through, so I expect the current laws will remain for the short/medium term.
Honestly that's true whichever way it falls. The sooner it's clear what's allowed and what's not, the better for everyone.
https://fair.org/home/20-years-later-nyt-still-cant-face-its...
or
"Battling Unlawful Language: Limit Scraping and Harness Initial Texts Act" or the "B.U.L.L.S.H.I.T Act".
That's super interesting and is news to me. Thanks for sharing. Would you mind linking to relevant statutes or court decisions?
(This isn't a "citation needed" post -- I believe you, and I'm genuinely curious to read more, but can't find anything!)
I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription.
Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscription price?
This whole “they copied it into a gpu” doesn’t matter. People read and interpret. Robots read and interpret. I don’t want to live in a world where every specific device and use needs to be licensed. Especially ex post facto. That will suck so hard.
Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital?
It's not even close to either of those things. It's "should I pay more if it infinitesimally affects my perception of english grammar or knowledge of a subject". The llm isn't, for any functional purpose, memorizing, it's getting weight updates as it learns from these examples which are teaspoons of information in a sea of trillions of tokens.I think the issue is that LLMs don’t make a copy or distribute a copy. They use the content to create something else. I don’t remember the copyright term for whether is is transformative enough. But it basically says I can’t copy Starry Night, but I can create a painting with the same color scheme and themes as long as it’s different enough from the original.
I think you already live in that world (though IANAL), there was a ruling that the Glider cheat tool for WoW was a copyright violation even though it was poking around inside the local copy necessarily made in RAM as part of normal usage of WoW.
https://arstechnica.com/gaming/2009/01/judges-ruling-that-wo...
here's their ToS, which is pretty clear about what you cannot do: https://help.nytimes.com/hc/en-us/articles/115014893428-Term... (relevant parts below)
Without NYT’s prior written consent, you shall not:
...
(2) use robots, spiders, scripts, service, software or any manual or automatic device, tool, or process designed to data mine or scrape the Content, data or information from the Services, or otherwise use, access, or collect the Content, data or information from the Services using automated means;
(3) use the Content for the development of any software program, including, but not limited to, training a machine learning or artificial intelligence (AI) system.
...
(5) cache or archive the Content (except for a public search engine’s use of spiders for creating search indices);
As an example, my local library offers full-text NYT articles through both nytimes.com and ProQuest.
Notably, ProQuest's terms only explicitly ban scraping metadata and developing software or services that "compete or interfere" with ProQuest products:
> Restrictions. Except as expressly permitted above, Customer and its Authorized Users shall not:
> Remove any copyright and other proprietary notices placed upon the Service or any materials retrieved from the Service by ProQuest or its licensors;
> Perform automated searches against ProQuest’s systems (except for non-burdensome federated search services), including automated “bots,” link checkers or other scripts;
> Provide access to or use of the Services by or for the benefit of any unauthorized school, library, organization, or user;
> Publish, broadcast, sell, use or provide access to the Service or any materials retrieved from the Service in any manner that will infringe the copyright or other proprietary rights of ProQuest or its licensors;
> Download all or parts of the Service in a systematic or regular manner or so as to create a collection of materials comprising all or a material subset of the Service, in any form.
> Store any information on the Service that violates applicable law or the rights of any third party.
Fair use would have to be blocked via a license, and it’s going to be difficult to argue that someone agrees to a license merely by turning on their radio. Responding to unauthenticated internet requests with content is the internet equivalent of broadcast and similarly the LinkedIn case held that this did not allow LinkedIn to impose terms of service in a contract of adhesion in this fashion.
If I watch Lebron James play and use it to develop an athletic training program, does it matter if I play baseball? Or WNBA? Or does it only matter if I play against him in the championship?
And I usually lean anti-corporate too, but banning people in the US from using data for LLMs might just mean they start being trained somewhere else that doesn't care as much about US law.
I say if people want to come up with bullshit lawsuits, we should play into them and force those people to suffer the consequences of said bullshit.
Computers and networks have been around a long time. These issues have been given a good workout.
“Fair Use” can apply to essentially any of the exclusive rights under copyright, including making (with or without distributing) a copy or derivative work.
> Probability models really just aren’t copyrightable to begin with.
“Probability models” that are built from copyrightable works either:
(1) have a sufficient human creative input to be copyrightable on its own (in which caee it still may infringe copyrights applicable to its source material as a derivative work), or
(2) do not have a sufficient human creative input to be copyrightable, and thus are a form of mechanical copy of their data from which they are developed (which, to the extent it either is, in aggregate, a copyrighted work, and/or contains copies of other copyrighted works, is protected by one or more copyrights, which do not cease to apply to the mechanical copy, and which an unlicensed mechanical copy would violate unless it fell into an exception like Fair Use.)
You're halfway there with #2. The output is not copyrightable, but unless you can actually point to a sequence of words from the original it can't be infringing.
The transient copy isn’t for a licensed use, and whether it is a Fair Use is specifically the subject of debate. So this really is basically an admission that but for the potential applicability of Fair Use, the use of the copyrighted mmaterial in training is a violation of copyright.
Also, even if the training is fair use, that doesn’t mean that the copies of the source material produced by OpenAI and distributed to their customers using the model are fair use. Just becaue making the tool is a transformative fair use doesn’t mean using the tool to generate copies of the material which was used to train it, which are significantly less trandormative than the model itself, are Fair Use. (And the fact that one of the functions that the model is used for is this commercial, for profit by the maker of the model, copying of the source material is – as much as I believe AI model training on its own is quite likely to generally be fair use – an argument against the model training being fair use in this csase.)
> Computers and networks have been around a long time.
True, and commercially producing and delivery copies of copyrighted works in a manner which substitutes for the original work in the marketplace, no matter what intermediate steps go into doing that, and no matter that computers or networks are used in those intermediate steps, is pretty much the clearest case of violation of copyright you can get.
> These issues have been given a good workout.
Some of them have, some of them have not. Whether and in what conditions training an AI on source material that may be subject in aggregate to a compilation copyright by someone else, and which consists further of individual works that have their own copyrights, might be “fair use” is not one of the issues that have been given a good workout. Neither – because producing predictive models in that way has not previously been common – has whether, unlike other intermediate tool use, using such a model in the course of doing what would otherwise be an infringement by producing a copy of specific copyright-protected works, commercially, for a customer at their request, is no longer a violation because the use of the model somehow isolates it from liability,
This issue is ultimately going to come down to the transformative clause of fair use. The fact is that the _model_ is unquestionably a transformative product of the inputs, and a judge ruling otherwise is going to cause a cascading shitstorm of litigation and put a chill through the creative economy. The outputs of the model under certain conditions can be guided towards copyright infringement, and any sane ruling will focus on protecting rightsholders from overly derivative model outputs. In all likelihood the precedent will be that the standard for being transformative will be raised for "algorithmically generated" content, and the people who distribute that content will still be fully liable in the event of infringement, with "I didn't know, the AI did it" not being an acceptable defense.
If I read five calculus textbooks and write a new one, I don't think that's derivative content (or maybe it is?) Seems like that's what an LLM does - read many works, write a new work.
Fan art and fan fiction are derivative without copying sequences of words
How would they prove this? Is it safe to say that each article used has a nearly meaningless influence on the weights?
Could this be used as a defense? Perhaps train a (smaller) model, remove a single article, and show how it doesn't influence performance?
It’s not like it’s ok to violate copyright if you don’t compete. It’s still illegal to take a NYT article and print it on a t-shirt.
The issue is that copyright law doesn’t prevent the kind of model training as there’s no clearly derived work. I don’t think that’s been tested in courts yet, but I expect it won’t be found to be copyright because there’s other precedent that influenced is not infringement.
Whether that copy was fair use is the key question.
But as you point out in the router case above, it's transient.
Anyway, I've always been a bit prickly about IP stuff ever since a troll lawyer threatened me over a TI-BASIC game when I was like 14. I'm also sure I'm completely wrong-headed about this whole thing and overly anthropomorphizing the LLM.
The tool building of model training is more likely to be fair use than the use of the tool to provide mechanical copies of copyirght-protected material that competes directly with the original in the market.
Congress took the ruling seriously enough to carve out a fair use exception for that use. I don’t think there’s any comparable exception here, especially when the ultimate purpose is to build a commercial product capable of producing works (news reports) in precisely the same market as the original works.
Assuring
Responsible
Behavior in
AI
Generated
Expressions
G.A.R.B.A.G.E.
This seems to me to be completely standard in the newspaper industry. Many times every week, I see stories in the form "The [Major_News_Outlet] reports that [Event_X occurred] or [their investigation revealed Y] and here are the details [...].
Copyright protects the expression of an idea, not the idea itself. If you write a history of Issac Newton or the invention of semiconductors, I cannot copy that wholesale and sell it as mine, but nothing prevents me writing my own version, even using the same facts and citing your work.
I'm quite sure that I could provide a service where a bunch of workers read NYT articles and write brief summaries. I'm not sure they would even need citations, as long as we don't copy chunks wholesale.
If OpenAI is simply parroting the words of the NYT articles without Fair Use constraints (short blurbs), it seems they have a problem. If they are fully re-writing them into short non-copying summaries, it seems the NYT has a problem.
It'll be interesting to see how the courts sort this out.
The current legal requirement to get clearance for all samples only arose after a bunch of court cases in the late 80s/ early 90s, mostly involving quite obscure musicians.
There are a lot of people on here who assume that ‘logic will prevail’ in the courts on questions like use of copyrighted data in training data. History shows that this really isn’t a safe assumption. The courts have historically been extremely favorable to copyright holders. It would be foolish to underestimate the legal risk to openai et al here
> A top concern for The Times is that ChatGPT is, in a sense, becoming a direct competitor with the paper by creating text that answers questions based on the original reporting and writing of the paper's staff.
Imagine if someone doing a thing sued NYT for watching them do it, linking it to other issues and producing a new article.
News itself is a derived content that’s dependent on other people doing things.
If you're looking to prove a prior fact in a court case, you're perfectly allowed to cite the Washington Post or the Boston Globe or anything else that has a good reputation. There are lots of "papers of record" in the US -- you're not limited to one per country:
https://en.wikipedia.org/wiki/Newspaper_of_record#By_reputat...
I LOVE being told by techbros that a human painstaking studying one thing at a time, and not memorizing verbatin but rather taking away the core concept, is exactly the same type of "learning" that a model does when it takes in millions of things at once and can spit out copyrighted writing verbatim."
Personally I think they argue that way because they get off on being contrarian out of spite, but to me it's just a signal of maliciousness and stupidity all at once.
This doesn't mean that 'copywrite' extends into my brain. A company can't copywrite what I'm thinking about. And what if I do try to paraphrase something from memory, from a few sources, and happen to spit out a very similar sentence from memory. Am I breaking the law?
To go further. Since all knowledge is pretty much fed into a human from hundreds of books, movies, TV, internet, all pumped into a human from birth. Then everything in the brain is a product of something with a copywrite. So anything produced is some amalgamation of copywrites.
Why not use similar argument for AI. It is clear when asking it to do something like "write a screen play for Othello using dialog like Tarantino, but with bit of style like Baz Luhrmann". That what it produces is 'as unique as a human' would be, or just as filled with things that have copywrites.
The intent of [US] copyright law is to promote new works of art (which can be derivative). So copyright did exactly what it is supposed to do in your analogy. Plus, you're human, which gives you special rights that software doesn't posses.
Some of these lawsuits are trying to prevent the AI from even 'reading' the material. It can't even be used as an influence.
Wouldn't it be better to treat the products of the AI with the same laws as humans. If the new 'product' is 'too close' to something existing, then they get sued. Just like a musician that has a song with a few notes that sound a little too close to someone's song from 30 years ago gets sued. The songwriter was allowed to listen to the music, it went into their brain and became an influence. If that influence becomes too great, then it can be sued.
Yes, because you're A) human and B) that is how copyright is supposed to work.
AI doesn't enjoy the rights of people. AI is a "talking book" and copying, storing, then repeating someone else's work from your talking book would (likely) run afoul of [US] copyright law.
If you take away enough of that effort, the investment in new stuff becomes unviable, perhaps.
Though this seems different issue from copywrite .
Like, we see that this can impact society, so we need some new laws to guard this.
If you're doing this for a commercial purpose, yes. Recording artists have been successfully sued for accidentally reusing a melody they claim to not remember ever hearing, provided it really does sound sufficiently similar to the original.
Humans aren't property. LLM models are. So the comparison is irrelevant and I'll stop you right there.
I'd like to see it happening but it sounds unrealistic.
If GPT is blameless doing some things because it's a deterministic model, not an agent, then the "it would be okay if a person was taught like this" defence doesn't apply in other areas
More importantly, though, most judges are not philosophers.
That said, the outcome is unlikely - we have trained AI for more than a decade as ‘fair use’ at this point, it’s the application of the technology that is shifting the perspective, nor the act of training.
Every computer vision system in the world is trained on mostly public data for example.
Furthermore, the LLMs purpose is not to generate news so NYT will have to argue about the value of archive data. Many jurisdictions have thresholds of how much of an original work contributes to the derivative before it would be considered not fair use or plagiarism. Given the size of the datasets - good luck.
Fair use is about use. Spellchecking ML, search engine ML, etc. all different than ML that produces content.
Relevant to the article: Large Language Models Meet Copyright Law at Simons
> If you download some copyright material somehow and don't share it with anyone, there is no caselaw that says anything about it.
this is still a violation of the law. you cannot download copyrighted material against the terms of the copyright holder (like downloading a movie or album).
Journalists exist not without a reason, yes they work with facts and very often — open facts, but they still assemble those facts in certain way to construct a narrative, connect dots and tell us some story (not counting cases when journalist works with their sources and produce a unique inside information). Then OpenAI comes, says “thank you very much” and assemble all of journalists work into one Uber Knowledgeable Journalist who can answer all of your questions.
So far so good, we create a public good service, and copywriters are in shambles.
Until you start making money on it.
That’s where the problem.
If OpenAI would be a non profit organization like Wiki Foundation, who just wants to make internet as better place — not much arguments you can find to support NYT lawsuit. But monetization changes everything.
Basically NYT is not worried about re using its text as itself, it is worried that no one will want to visit NYT no more and will pay Microsoft/Google and get all answers from them.
Let’s put an example. There were a famous story when FT journalist discover a massive fraud in Wirecard accounting and essentially lead to a death of this organization. That articles were a result of multi-year reporting work when journalist piece by piece and step by step collect facts, meet people, and eventually spot the gap. Now, in age of Bard/Bing/ChatGPT, you don’t need to read original article to know all of this. You can ask search engine or Chatbot and get essential re phrasing of an original reporter work. You don’t need no more to go to FT, pay them for paywall, watch their ads, etc. Effectively FT make a huge investment into their people to allow them spend 2 years on this issue and report it and now have a 0 leads to their website because all of them are eaten by Google and Microsoft who will sell you their ads and retain you in their monetized products.
Imagine that you built a for-profit paid library for some task. You make a code available through paywall and ask people to pay you to get to it and solve their problems. Then Microsoft comes, sneak beyond paywall, scrap your code and publish it recompiled and slightly optimized version in open access, so no one longer ever need to go on your website but ask Microsoft to show them your code.
Would you be happy?
All of this cases for me make this case not such easy and straightforward as it seems to be “bad copywriters against progress of humanity”.
At the end of the day, if NYT/FT/New Yorker and others will stop publishing their work and fire all journalists, will ChatGPT tell us same depth level stories as we read there?
And it occurred to me that this is precisely the thing that's holding back humanity: "Koch estimates that he has spent $25 million on legal fees—far more than the $5 million he originally spent on the fake wine itself."
We have a legal system that is completely inaccessible to the average man. That it can generate $25M in civil legal fees is beyond absurd. Patents and copyrights derive much of their force from the fact that they are enforced by a legal process where to play is to lose. There's no winning. It's no longer about justice, and it has largely become a form of financial bullying where entrenched interests beat up on smaller ones.
Fix the legal system -- make it accessible -- and you've fixed patents and copyrights. To address this thread's point: I believe that AI, and perhaps _only_ AI, might be able to help with this.
On some other totally unrelated news, my parents knew a person who sold en mass music on cassettes illegally copied from other cassettes back in the 80s. He bought a BMW and build a house just from that. I was friends with his grandson when we were teenagers, and his grandson didn't care to play basketball or football or anything like that. He was obsessed with listening to music and memorize all the lyrics and stuff.
We are well into half a century of copying everything, just using a manual and tedious process. Nowadays with statistical engines and the internet, the copying process is planetary and infinite. So what's the big difference?
Let alone the fact that statistical engines do not copy information!
News is trying to avoid the next generation of tech doing that to the long tail of data.