(apologies for the run-on sentence - it is early still)
(apologies for the run-on sentence - it is early still)
Technically the user is the one misbehaving, but we, Facebook, and any reasonable court know that users are doing that.
Computers cannot create copyright. They are not creative. Just because I save your image as webp or jpeg or whatever doesn't mean I have changed the copyright. Just because I zip it up with a hundred other images doesn't mean the zipfile is free of your copyright.
Effectively, computers are executing math, and math by itself does not construct new copyright, since copyright is the result of a creative human process.
As far as I can tell, current AI are fundamentally not too different from wildly complex compression algorithms. You compress a billion images down to a model. The model now can reproduce a fraction or the whole of the copyrighted work with some low probability. Rote and probabilistic compression.
The creator of the AI might own the copyright for what it produces if constructing the AI was suitably creative, i.e. if you construct an AI that trains on random noise and produces images, those are clearly something you, the author of the AI's code, can claim copyright over... But current AI seem like math more than anything else. It's plausible that reinforced learning or some other part of training does imbue creativity into the process, but that doesn't seem obviously true to me.
Also, saying that it's math ergo it's not creative is something that most people on HN would not agree with.
As for "it's just compression" - compression means that you can recover the original data - perhaps with a loss of quality, but still you can. With modern ML you mostly can't.
By the mere act of using Photoshop, no. By the act of providing your own inputs to Photoshop, yes.
However, when using the current image gen AIs, the input you provide is a sentence of text and a couple parameters, a minimal amount of creativity.
This would be akin to opening photoshop and doing minimal work, such as choosing "resize image, apply blur filter".
If you open photoshop and do a few rote transformations, you indeed have not imbued enough creativity to create a new copyrighted work, the work retains its original copyright if you just open it in photoshop and resize it.
Have you tried creating art with AI? Usually it takes hundreds of iterations of text-to-image, image-to-image, inpainting, outpainting using dozens of different models.
"A sentence is all it takes" is like saying all it takes to make a million is crossing some numbers on a grid.
See eg https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
308.2 Creativity A work of authorship must possess “some minimal degree of creativity” to sustain a copyright claim. Feist, 499 U.S. at 358, 362 (citation omitted). “[T]he requisite level of creativity is extremely low.” Even a “slight amount” of creative expression will suffice. “The vast majority of works make the grade quite easily, as they possess some creative spark, ‘no matter how crude, humble or obvious it might be.’” Id. at 346 (citation omitted).
The VAE can be thought of as a codec, but the denoising process can recover images that are far removed from anything that is in the training data. Nobody has ever created an impressionist painting of Winston Churchill riding a purple lizard through the gates of retrofuturist Constantinople, yet almost infinite variations of that image exist in the latent space. If anything, it can be thought of as an intricate form of collage, which we do give special treatment for copyright purposes.
They store the image or video (host/copy), distribute it over their network and to users (use/run), they resize it and change the image format (modify/translate), their site then shows it to the user (display/derivative work), and they can't control the setting in which a user might choose to pull up an image they have access to (the "publically" caveat)
It sounds like a lot, but AFAIK that's what that clause covers and why it's necessary for any site like them.
IANAL, and my jargon may be off, but I think that in the scenario where you get sued for something that's been litigated to fall under this clause in the past, you can basically say "even if we assume the evidence and claims are accurate, it's obviously in the clear based on prior cases", if the judge agrees, you win without going to trial, which is a "summary judgement" I think.
On the flip side, if someone is trying to apply the clause in a novel, not previously litigated way, you're way less likely to get that summary judgement and it will have to be argued in court.
It works the other way too, if I wrote a eula that used different phrasing than what's been established prior, say to make it more obviously cover just the normal stuff for user uploaded images, summary judgement is less likely to succeed because no court had ever weighed in on my novel phrasing as covering those actions in that way.
There's also the risk that if you make the phrasing too narrow (specifying resizing of the image) then when a new tech comes along that's reasonable to apply (e.g. some ML process to derive a 3d scene from images, or make them) exactly zero user uploaded images you store at that point could benefit from that until you go back and ask the user to agree to that too. The question then becomes how worth is narrowing the wording when you can accidentally paint yourself into a corner.
Or how about if it had been phrased "display on a monitor" had been used years back pre-smartphone era? You could be sued for making user uploaded media available to view on phones since that wasn't in the license granted to you by your users!
When you cover all the little edge cases, you end up with the seemingly overbroad clause most companies use.
An important thing to remember is that the legal interpretation of a text can differ almost arbitrarily from the plain English meaning of the text as written.
Well .. no. It happens each time that Google et. al find a new way to use your data. It's what all we German "privacy nuts" have warned people about for years and the reason that the older German data protection laws and now EU regulations require you to state exactly what you are doing with data ("purpose limitation"). If companies can just write "oh well, we will use it for something" how can anyone evaluate whether they should accept without knowing the future? Right. They cant.
So, this could be another case of the EU kicking Facebook in the face. We'll see.
The problem here is cowboyscott doesn't own a copyright of Ironman image. But his uploading of image may match the condition of fair use of US or similar copyright exemption rule in their country's copyright law. It effectively works as copyright laundering.
But I don't know if it's really laundered anything. If you say "Hey Meta AI, make me poster for my cookie company that has Iron man eating my cookies" I'm pretty sure Disney could still sue you. It could still sue you if you instructed a human to draw a picture that had Ironman in it so I don't even know if you need a new legal framework.
DMCA take downs seem to feel that this is not a thing any longer.
Personally, I'm on a side of using copyrighted data for machine learning input source doesn't violate copyright. Statistically, learned model for generative Ai doesn't retain even 1 bit of input. It's hard to say NN model data infringe any copyright of the input source. The copyright is applied to the expression, not the process. If the generative AI produces an image that's clearly a copy of a specific Ironman image which existed before the image generation, that's copyright infringement.
[0] https://stackdiary.com/chatgpts-training-data-can-be-exposed...
Obviously not every generative output is a copyright violation, but it seems equally clear that there are outputs that would be if they were produced by humans.
It does. The data is just obfuscated.
> When you share, post, or upload content that is covered by intellectual property rights (like photos or videos) on or in connection with our Service, you hereby grant to us a non-exclusive, royalty-free, transferable, sub-licensable, worldwide license to host, use, distribute, modify, run, copy, publicly perform or display, translate, and create derivative works of your content (consistent with your privacy and application settings). This license will end when your content is deleted from our systems. You can delete content individually or all at once by deleting your account.How derived data is handled after copyright is revoked is a question thats hard to answer.
I suspect that the data will be deleted from the dataset, and any new models will not contain derivatives from that image.
How legal that is, is expensive to find out. I suspect you'd need to prove that your image had been used, and that it's use contradicts the license that was granted. It would take a lot of lawyer and court time to find out. (I'm not a lawyer, so there might already be case history here. I'm just a systadmin who's looking after datasets. )
postscript: something something GDPR. There are rules about processed data, but I can't remember the specifics. There are caveats about "reasonable"
Huh? I think you want s/(:?m[^m]*)m/tr/
But in the US this hasn't been tested in the courts yet, and there's reason to think from precedent this legal argument might not hold (https://www.youtube.com/watch?v=G08hY8dSrUY - sorry don't have a written version of this).
And the lawsuits so far aren't fairing well for those who think training should require having copyright (https://www.hollywoodreporter.com/business/business-news/sar...)
As well as learning, as a whole.
Unless there is literally a substantial copy of some particular piece of copyrighted material, it seems to be a massive hurdle to prove that analyzing something is copyright infringement.
If anything, I think that severely hinders the pro-AI argument if fanfiction made by human authors are also bound by copyright.
ETA: I just tested it out and you can totally create Interview with a Vampire fanfiction with Bing Compose. That presumably is subject to at least as strong copyright as human authors and is thus a copyright violation.
> Copyright protection is available to the creators of a range of works including literary, musical, dramatic and artistic works. Recognition of fictional characters as works eligible for copyright protection has come about with the understanding that characters can be separated from the original works they were embodied in and acquire a new life by featuring in subsequent works.
Creating a work using Harry Potter or Darth Vader or Tarzan ("As of 2023, the first ten books, through Tarzan and the Ant Men, are in the public domain worldwide. The later works are still under copyright in the United States.") is a copyright infringement.
You may also find https://www.hollywoodreporter.com/business/business-news/dc-... interesting as well as the entire legal saga of Eleanor.
---
Creating Interview with a Vampire fan fiction with Bing - Bing didn't have any agency. The question of copyright infringement (I believe) should be only applied to entities with agency to (or not) ask for copyright infringing works.
Transformative works are a thing:
https://www.transformativeworks.org/faq/#:~:text=investments...
https://www.transformativeworks.org/faq/#:~:text=Open%20Door...
That’s the output of the model, it doesn’t have much bearing on the copyright status of the model.
Satire, criticism, reviews and journalism are explicitly permitted under fair use.
If I wish to publicly express my disdain or praise for your art, it is necessary that I can show samples / pictures/ photos when I express whatever my deal is.
The portion of the training set might. The actual trained result -- the outcome of a use under the license -- would, at least arguably, not.
Of course, that's also before the whole "training is fair use and doesn't require a license" issue is considered, which if it is correct renders the entire issue moot -- in that case, using anything you have access to for training, irrespective of license, is fine.
The interesting question is just who will be liable for the copyright violation: The party that hosts the AI service? The party that trained it on copyrighted images? The user entering a prompt? The (possibly different) user publishing the resulting image?
When Disney did their copyright extension last time, they had bipartisan influence.
Now Disney is in the middle of the culture war, and there is no Republican that will risk being primaried to support Disney.
Given that you de facto need 60 votes in the Senate, it is not happening.
the same is true for artwork.
Is selling colored pencils that can draw images of Mickey a copyright violation?
The way I see it, the tool can't ever be at fault for its use, unless its sole use (or something close enough to its sole use) is to infringe in copyright.
Besides, the safeguarding of copyright isn't the single variable we as a society should be solving for. General global productivity is way more valuable than guaranteeing Disney's bottom line.
Even then, you could look at a tape recorder or a photocopier and one of their primary uses is to make a copy of a copyrighted work.
The question isn't "can it be used for" but rather "does it have valid non-infringing use" and "when it does infringe, is it the person who uses the tool or the tool that is at fault?"
That is certainly not clear, unless its only purpose was to do that.
I don't think that courts have ruled on that specifically (yet), but I seriously doubt that it would be. Taking the image of Mickey and distributing it would certainly be, though.
A photocopier can generate images of Mickey. Does that make a photocopier illegal?
That stance is clearly not supported by copyright law.
If, however, we're talking about copyright violations applying to the distribution of works generated by AI, that's an entirely different conversation. It's still not really clear-cut, but there are ways that could be in violation of copyright law.
It isn't the case that AI is being treated differently, though. The issues would be the same if a human were doing all of this stuff.
Are tattoo artists breaking the law by creating tattoos of copyrighted material? I think they are. And if an artist becomes really popular for their mickey mouse tattoos, then they will provably be noticed by Disney and there will be consequences.
I think AI companies are working hard on preventing generated images from being similar to training images unless the user very explicitly asks the result to look like some well known image/character.
I don't think this is going to be hard for courts. If you borrow your friends copy of a copyright text, got to kinkos and duplicate it, then distribute the results - you are the one violating copyright, not your friend or kinkos.
The same will hold here I think, mutatis mutandis. This is all completely separable from the training issue.
This isn't about copyright, it is about the fact that most people don't realize that by posting photos, they are licensing those photos.
If they were scanning my private messages, things would be different.
1 - human experience ends up informing human ingenuity. A sketch of Wile E. Coyote comes from someone’s (Chuck Jones?) experience of dogs and seeing coyotes, plus innumerable experience with things that are funny, constraints from experience of certain features that do or don’t work well on animation cels etc. Perhaps a stray tweak in his ears come from a Rembrandt seen as a child or from a glance at a sketch in progress by the person sitting at the next easel in a drawing class long ago.
In todays’s jargon our experiences are all parts of our training set (though today’s massive RNN models are infinitesimal by comparison).
And I think of my tools the same: a ton of inputs stirred together is fine by me.
2 - a difference is that fb’s model is made from public posts: posts offered for anyone to see. In the human case even my private experiences are part of my “training set.”
However, as others note, all the actions of the major IT companies indicate that their legal departments feel safe in assuming that training a ML model is not a derivative work of the training data, they are willing to defend that stance in court, and expect to win.
Like, if their lawyers wouldn't be sure, they'd definitely advise the management not to do it (explicitly, in writing, to cover their arses), and if executives want to take on large risks despite such legal warning, they'd do that only after getting confirmation from board and shareholders (explicitly, in writing, to avoid major personal liability), and for publicly traded companies the shareholders equals the public, so they'd all be writing about these legal risks in all caps in every public company report to shareholders.
I think the move will be to argue fair use, declaring the derivative work to be transformative, and possibly to point out that only a small amount (1%-3%) of the original data is retained.
I don’t think there’s a question that people are allowed to upload photos like that.
Technically, that's a copyright violation. Disney just opts not to enforce their rights for that sort of use.
Similarly, you technically can't take and post pictures of statues, paintings, some buildings, etc., and some rightsholders do enforce their copyright when people do those things.
Outside of things within the scope of fair use would be within Disney’s rights to restrict, but given the actual public policies and guidance on photography at Disney parks, I think there is a very strong case that noncommercial photography (for people are present as paid guests) is permitted by implied license.
> Similarly, you technically can't take and post pictures of statues, paintings, some buildings, etc., and some rightsholders do enforce their copyright when people do those things.
Well, not buildings if they are in or visible from a public place in the US, at least under copyright law. (Photography of some, particularly government, buildings may run afoul of other law.) This may be different in other countries.
Ahhh, you're correct. This was apparently changed in 1990. I just hadn't updated my mental model in accordance with that change.
https://www.nolo.com/legal-encyclopedia/copyright-architectu...
If you can[0]crawl materials from other sites, why can't you crawl from your own site?
[0]: "can" in quotes
When new mediums are invented (like internet), you need to sign an annex to the agreement extending it to this medium.
Having said that, I would still consider it a fair use to train model on given images, but using the trained model to replicate a specific style etc, would most likely be considered a new medium. (IANAL though)
Doubt it. If you upload child porn to Instagram and they distribute it - it's still an Instagram problem, AFAIK.