Japan Goes All In: Copyright Doesn't Apply to AI Training
biia.com
biia.com
A potential downside is that AI systems can 'mechanise' the creation of material that potentially infringes copyright (in the same way that human generated content can infringe)
But a potential upside is that we can 'mechanise' the process by which we judge whether new content infringes the copyright of older material.
If we don't take this approach then there will be a series of very lame legal loop-holes with putting mechanical Turks [1] in the process. Or just end up with very "I know it when I see it" legislation.
So for both practical and philosophical grounds I do support this.
The end is inevitable, protecting copyright for AI training is probably a lost cause.
And while I expect NYT imagines that their archive is really valuable for training, they're just not that special in the sense that while they may have broken more stories on average than many others, and have had influential op eds etc., the ones that matters will have been cited and referenced and written about elsewhere - the irony is that by virtue of being so well known, their historically most important content is also less unique in terms of the accessibility of the information in it.
So while I'm sure OpenAI would love their archives, I'm also sure that if OpenAI and others have to license content and NYT end up being "difficult", OpenAI will just license content from (or buy) a suitably diverse portfolio of other papers instead.
In other words, beyond producing outright synthetic data, if AI companies are prevented from training on data they don't have a license to, the net effect will just be a scramble to buy licenses and/or buy companies that can provide sources of content, and the price for that content will be a lot lower than some of the people pursuing these copyright claims imagine.
In the end, if we go that route, all we'll have achieved as a society is creating massive moats protecting the companies already big enough to buy access to a broad enough set of content and made open models harder.
Exhibit J (phone won’t cooperate with paste).
Generative AI algorithmically processes the inputs such that they can be roughly recreated after decoding. A bit like how jpeg encodes images "beyond recognition"(aka lossy).
[0] https://www.technollama.co.uk/high-court-rules-that-getty-v-...
[1] https://www.theverge.com/2023/1/17/23558516/ai-art-copyright...
The same thing happens for images - the more an output looks like an input, the more likely (as diffusion-based generation proceeds) it is to look more like it. There are recent examples of Dune posters being recreated essentially as-is.
The issue is that generative AI - both for images as well as text - in effect memorize sources as well as learn from from, and can end up regenerating training sources verbatim (or with minimal changes in case of images).
I don't think any US court is going to accept "yes your honor, we copied this copyright material, but we used a TOOL to do it" as a way to avoid copyright.
I believe human artists consciously adjust their output to avoid copying previous artists too closely. And sometimes they choose to copy very closely or exactly.
Obviously the same feature can be implemented as an option on generative AI systems.
No doubt generative AI systems could be built to self-police and not emit any tentative outputs that are too close to training samples, but that's certainly not the way they are today, and it's not clear from the article that started this topic that this issue is addressed in any way by the Japanese law, which is just about training data.
An AI training itself on a million newspaper articles can be declared legal, sure, but what happens when it also starts spitting out the same articles with nearly no modification? Is that still fair use? This is the crux of NYT's lawsuit, and making laws about AI training isn't going to make a difference to that.
The claims of NYT are more than just about training. They're claiming that as part of the ChatGPT software; it _looks up_ stuff in a database of articles. That is beyond _training_.
If a human kept around a briefcase of NYT articles they didn't pay for and let you view them for a fee I think everybody would agree that's copyright infringement.
The focus on "whats the difference between an AI and a Human learning" is a great slight of hand to bypass the obvious copyright infringement OpenAI is doing.
So essentially you can ask chatgpt to fetch and summarize a paywalled article, because openai has subscribed their crawler?
Specifically the NYT put in the first sentence of the article and asked GPT-4 to autocomplete it, which it did with >95% accuracy. It's really quite stark: https://nitter.net/jason_kint/status/1740146134767865895#m
The issue isn't that GPT-4 users can read NYT stories without paying for a subscription, though that is a legitimate concern. The issue is that for good-faith use cases - e.g. asking to write a summary about a recent current event - GPT-4 could very well copy an entire paragraph from NYT verbatim, without the user having any way of knowing. It's a serious problem.
You might be thinking of newspapers’ political campaigns such as C-18 in Canada, which were done in bad faith. The newspapers wanted links to stay but wanted money too.
OpenAI can’t do this for NYT without destroying their model and remaking it.
Except that's not what it's doing. Show me how I can get ChatGPT to show me the full text of a NYT article.
https://www.courtlistener.com/docket/68117049/the-new-york-t...
Did ChatGPT change this after the lawsuit was filed? Probably. Does it matter to this conversation when it's clearly possible to limit verbatim outputs of copyrighted text? No.
There are like 40 pages of examples in NYT's lawsuit showing exactly that.
Sure. Slightly more interesting is if that same human with those same breifcases was taking money to answer questions and referenced those papers, but did not just provide the article or headlines, and might not even be paraphrasing the article at all. Is that okay?
To the extent it is merely paraphrasing articles, or outputting headlines that it just looked up, I agree that could well be infringement. If it more transformative processes those articles into something distinct, then it is not nearly as clear cut. The latter is arguably the intent of openAI, even if the current results might be closer top the former.
That is clearly not the case. There isn't anything like enough storage space in typical LLMs to maintain a "database" of all the training data.
If an article can be reproduced from an extremely sparse representation using a stochastic algorithm, that seems to me to be prima facie evidence that the article didn't contain much (or possibly, any) significant creative content to begin with.
Certainly I would expect to find less creativity in an allegedly factual news article than in a piece of acknowledged fiction.
Only creative works are copyrightable in the United States.
There are soi-disant "artists" who produce "artworks" that are (e.g.) nothing but a pure white rectangle. That doesn't make <div style="background-color:white;width:100%;height:100%;"></div> any kind of copyright infringement.
I don't see how it is obvious, nor how it is infringement.
How come this doesn't apply to the person who memorized the NYT article, and recited it verbatim?
It's obviously infringement for the person who pressed the "generate" button to produce the article. That's no different from someone who copied a picture using photoshop. However, photoshop itself (and the making of it) does not constitute any infringement whatsoever, as long as at the time of making the application, the sources used are not infringing (which, i presume openAI had the right to view the articles at the time of training).
The crux, to me, is that the information extracted and produced (aka, the neural weights) do not itself constitute any infringement. Using those weights to generate copyrighted stuff is an infringement, but only for the person _doing_ the generation, not on the authors of the weights.
How much modification is enough modification? The courts can wrangle with that endlessly.
If AI training maximally wins, it could substantially erode the value of IP to the degree of threatening the business models of some very important societal pillars like news media.
That would not be transformative. If it is also provided as a commercial service that directly competes against the copyright holder, then that would not be 'fair use'.
All kinds of interesting legal questions.
However system will absolutely destroy you if you try to sell or distribute copies of material you collected...
But here in Mexico, you can download all you want, and last time I read the law, as long as you did stuff "not for profit" you could also distribute digital content.
(of course I always say that, here in Mexico it is illegal to murder people, kidnap and whatnot, and look at how much they prosecute people that do that [95% of crime goes unpunished in Mexico [1]])
[1] https://www.nbcnews.com/news/latino/violent-crimes-rise-mexi...
As I said in another thread: Sorry if your 40 hour work won't pay you $10 bucks a month forever. That's the case for most of the rest of us: we produce for 40 hours, we get paid for those 40 hours, regardless of what we do. Welcome to the club!!
Exactly. It would be entirely legitimate to legally privilege human learning without allowing that privilege to transfer (by analogy) to machine learning.
There are a lot of people who want use analogy to force the transfer of that privilege because 1) they judge they can profit handsomely and/or 2) they're technology enthusiasts who've read too much sci-fi.
Specific IP like characters are already perfectly well protected by copyright. It does not infringe on anything to learn about Batman.
Just “training for the heck of it” without any justification for limits besides hurting the artists feelings who put their art in public kinda sucks for them but the alternative is just saying that only but doing otherwise is just asking for only large corporations to be able to train. And it’s not like artists will be comped for that either. It’ll just be them losing out because they put it on a “free” platform.
Because if you think this is not ok then you are arguing that people own styles. And they don’t.
And if you think it is ok then you’re just arguing for pointless extra steps.
I’m perfectly within my rights to make art with your artistic style. And I’m perfectly within my rights to train a bot on my own works. So what’s your argument? What am I not allowed to do with this poorly conceived law?
"the art" -> the art you made that you have rights to use, or "the art" that someone else made that you don't have rights to use?
>It’s just a matter of whether you needed a human in the middle
No, your story was you made totally new art that you had rights to use to train your AI. It's not a human in the middle, its a human author who allows you to use the art at all. If the "human in the middle" in round two didn't give you the rights you couldn't use those either. The human is doing the authorship of the work and also allowing or not allowing you to use it. They aren't in the middle, they're 100% of the issue and the difference between allowed and not allowed, human authored or not human authored.
It’s ok to make a LegitShady bot as long as you can pay someone $30 to make a few works stylistically similar to yours?
Because if that’s true, you’ve just agreed that the value of your creativity is $0. You don’t even get the $30 here. Nobody really wants your specific works, and your creativity will be freely available. In fact it can probably be synthesized with just a human and some simple tooling in the near future / now.
Is that what you want?
These are the consequences of your own rhetoric. You are ignoring them for cheap shots to take the moral high ground. You can’t just say “I’m the side of artists” and conclude that you’re doing good. Your rule set is VERY WEAK and prone to abuse. It will not meaningfully protect artists at all.
The copyright claims are spurious, the laws were never written with such powerful algorithms in mind. If copyright applies to training, and if law is based upon principles instead of raw power then such a ruling would lead to strange places.
Bravo.
This to me seems like the right approach, and is not much different from what humans do. Humans are free to read whatever source material they want, but you can't subsequently write it down from memory verbatim or nearly verbatim.
edit: And if you think I'm being hypothetical, Github Copilot already does exactly this.
The only possible way to do this is for companies to provide a list of all of their sources and for me to then automate verification and hope it works!
The real answer is that I should be able to control whether my information is used to train models or not, because we already know that models spit out verbatim results with generic queries and there’s just no way for a user to otherwise check this.
Regarding you wanting to opt-out of training, that's fine you can already do that for many large models. But likely doing that will become the equivalent of putting your works in a safe where no one but you will ever end up reading them or finding them.
Banning all AI training is the equivalent to banning search with any modern search engine.
edit: Also, note that you are already on the hook for not infringing patents, which if you were to try to do "perfectly" like you imply with copyright, then you need to search/read/understand the entire body of published patents. A task that is clearly not possible. Yet patent law functions (unfortunately).
What good is training an AI if nobody can use it?
Pretending that disallowing AI is the same as disallowing everyone is a strawman. People finding you is a whole lot different than your work being copied being completely okay so long as it was an AI that did it.
... or have a computer "type in" a book i.e. file copy i.e. OCR scan (or microphone for audio) (or video-record for video) ...
It's no different than a photocopier or a VCR when you're using it for that purpose, it's the end result that matters.
Time will tell.
Interesting times!
We'll see how Sony Music Japan, Universal Music Group Japan and the rest think about this latest iteration of regulatory arbitrage with AI.
[0] https://stability.ai/research/stable-audio-efficient-timing-...
South Korea has mirrored Japan's policy: https://metanews.com/south-korean-government-says-no-copyrig...
Obviously, there is a difference between training AI and using AI.
Its very hard to make a case that training AI breaches copyright in any way.
But its copyright 101 that if what that AI spits out after isn't transformative the AI retailer is getting hefty fines.
(n.b. this is a complex problem and the scenario isn't quite so simple, if I hire an artist that creates work with a copyrighted figure / use a site that gives me random images it created that I know to sometimes replicate copyrighted figures, its not the artist / site / even me having the image in my possession that's a problem)
And at any rate, a country banning this tech will be missing the revolution. Protecting the buggy whip manufacturers and all that.
Speed and scale make these completely unrelated in my opinion.
I definitely agree with this, but I think a "good analogy" is really rare.
In this case, I think that the analogy just obscures the main point about copyright. I think the original comment would have been stronger by omitting the analogy and just discussing where the legal responsibility falls.
In almost every case an analogy is made, people discuss the validity of the analogy over the actual point of the comment (I'm guilty here too! my comment ended up being about the analogy rather than the main point).
While the ideals of freedom these brought may be desirable, we may want to avoid repeating the bloodshed from this particular bit of history. And that goes double for anyone in government who takes personal exception to being treated like French royalty.
Never mind that every regulation that constrains AI development in the West is a giftwrapped blessing to China and other actors that DGAF about copyright law.
You'll make friends on both sides of the political aisle with that position, but it's a shame to see it so readily accepted around here.
(Can't reply due to HN's rate limiting algorithm that penalizes me after four or five posts while allowing the most-corrosive trolls imaginable to party all day, so I edited to clarify.)
What?
The difference as far as copyright is concerned if one person has a scanner, camera, or blank cassette tape or everyone has them are the same.
If telling AI to study Spiderman and then output 10 pictures of Spiderman is illegal, how is that different from hiring 10,000 artists to study Spiderman, having them each do a drawing, and then hiring 100 talent judges to pick the top 10?
I think the more immediate issue with AI is that it's like having access to a close-to-zero cost human who doesn't care whether or not they're creating content which, if a human did it, would be considered a copyright infringement. And that they care so little about copyright (and other data rights) that they're basically incapable of even warning you if they are close to an existing character or living person.
I don't know how this is going to play out, but right now we're getting a more polite re-run of the luddites smashing early industrial equipment. I can sympathise with the loss of purpose and economic disenfranchisement, but the economic power in that revolution went to those that did the most automation, and I expect the same to be true this time.
Just because I drew Mickey Mouse doesn't mean I can sell it. Just because I sing Taylor's song doesn't mean I can upload it to Spotify.
Just because GPT can return an image doesn't mean I'm allowed to sell it.
Creation and distribution are different, and the reality is that the 90% of consumers will never create, they will consume distribution.
This is a very weak argument.
A pencil, an empty USB stick, an LLM without weights, and a small child can all be used to reproduce copyrighted works. A USB stick with a copy of a movie on it, an LLM trained on a book, and a human that has watched a Disney film can all actually infringe copyright.
There's an issue beyond copyright, though. A lot of companies seem to think it's okay to train LLMs on data that they at least have no moral rights to and possibly have no legal rights to either. (A TOS saying that anyone posting anything privately, IMO, does not mean that the person posting it had rights to it, nor do I believe that fine print ought to give anyone rights to anyone else's private information.) And a person that read all your email and an LLM that has been trained on all your email can both easily infringe your rights to privacy.
I find this a pretty week argument too. Does an LLM use a pencil? Can an LLM use a pencil? Does a pencil watch anime or read the NYT? Does a pencil create new work when upon request?
Are these things really comparable? Are they even related?
TLDR: A "Moloch Trap" is a generic term for situations like the the Prisoner's Dilemma, where optimal actions for the group are the opposite of optimal actions for each individual in that group.
Moloch didn't win here. Everyone else did. We get new technologies, the copyright owners got everything they were promise, just not more.
I agree that the result is not a "Moloch trap": as you say, the new technology actually does de-fang Google Search and empowers many smaller people to be able to do things they would never otherwise have been able to do. The contrary ruling would certainly have been enjoyed by "copyright maximalists" like Disney extracting value from the public domain without giving anything back.
But there are more people affected than just greedy copyright maximalists. Individual artists who have spent years developing a distinctive style and making it popular are seeing their style copied ad-infinitum for free. Organizations like the NYT that invest money doing investigative journalism are having their results slurped up and regurgitated.
To your comment:
> If a person is allowed to read training material acquired legally, then so too must their LLM experiment be allowed to read it. The LLM reading it creates no new copies.
In the past, each copyrighted work seen might train a single BNN (biological neural network). Only a small percentage of BNNs would actually study such work to learn to emulate it; only a handful would achieve parity or exceed the quality of the work. Each BNN was expensive to employ, would only work for a certain number of hours per day, and a certain number of years before retiring.
Now a single ANN (artificial neural network) can study works to emulate them in a month or two. That ANN is far less expensive to employ than a BNN; can be deployed 24/7 indefinitely; and can be duplicated to as many GPUs as someone can get their hands on.
Currently legally, it may be that an ANN learning from an artist's work is the same as a BNN learning from an artist's work. But from a practical perspective, from the case of individual artists, it's clearly not the same.
Now maybe that's the inevitable price of progress; but 1) I don't think that's the inevitable conclusion, and 2) even if it is, we need to be honest about it.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
So if i just zip up copyrighted images using a NN, then, what? They're public domain?
Regulators here are miles away from understanding the implications -- this is what happens when you let companies whose profit motive is selling "AI" be the "Experts" on the topic.
If you don't care about Disney, fine -- so what about your health records? This is also the prelude to an end to privacy
Great. So d'you think you could outline a reason why you wouldnt have an interest in your creative works not being used also?
Either the training data is, as big-ad-tech says, essentially equivalent to generic human experiences -- ie., weakly repoducible; OR it is extremely reporducible, and equivalent more to standard contemporary data compression.
If you're kool-aid'ing the former on copyright, why not the latter on privacy?>
So arguments like "I can get the AI to output my chart verbatim" start carrying weight because it's granted access to data that the humans that created the AI are not permitted to share in any form whatsoever where as copyright concerns what I may do with the data after it's produced. Copyright is full of exceptions for things that don't count as a reproduction or performance of the work and this is just one more, it doesn't change the nature of copyright.
> There is no difference between a lossy jpg, taking its pixels as weights, and the weights of a NN.
Somehow I don't think this is going to hold up in court.
Can you help me understand the relationship here?
Something can be protected by privacy laws without being under copyright, and vice versa.
W.r.t to the NYT case - It's my opinion that it's completely reasonable to use a corpus of vetted english literature like the NYT as a way to train your model to comprehend language - but if the model also begins to echo the contents of those articles then that may be a serious breech of the NYT's right's to monetize their work.
"zip up copyrighted images using a nn" is trek level technobabble.
that's not how NNs work.
how the hell does it even connect to privacy?
copyright isn't what makes it illegal to expose and have your medical records
it's privacy violations, which this doesn't even touch
Look up 'overfitting', neural-network based compression, etc. or that paper that used zip compression as a neural-network basically. Farthest thing possible from being 'technobabble' once you understand how inextricably linked compression and 'understanding' is.