I've resigned from my role leading the Audio team at Stability AI
twitter.com
twitter.com
Maybe current AI output is still too similar to the training data, but that seems like more of a reason to regulate the output and not the input. We already have legal frameworks to prevent people from replicating the work of others. I don’t understand why there needs to be a distinction from using copyrighted works to train computers vs people.
The problem isn't that the works are copyrighted. It's perfectly ethical to build a business on licensed works in order to generate new text, imagery, audio, etc. based on those works. Adobe Firefly is an example of ethical generative AI trained on imagery Adobe owns, licenses, or is in the public domain.
As opposed to a “free for all” where they get literally nothing?
Are we really arguing the morality of who gets to screw over artists?
:|
My sympathy for people who want to be a part of that screwing them over rather than letting big companies get all the glory of screwing them over isn’t really very high.
Human capacity is rate-limited, and that’s an important distinction.
Another point: humans exist for a finite time. If I learn how to write music per the above, I’ve only got a set period of time wherein I’m collecting from the talents given to me by the shoulders upon which I’m standing. Sooner or later I retire/die and someone else must become the cultural torchbearer. If you consider some finite amount of capacity in the world to gainfully support all artists, I’m only drawing from that for a finite period of time.
Contrast that with AI, where, theoretically, in 50 years a model will still be just as able to create outputs as it is today.
Having said all that, I don’t have enough skin in the game to decide whether that builds an argument for regulating input vs output. I’m certainly sympathetic to the idea that it might, though. It’s a complex topic and saying “humans work this way, computers are fine to as well” doesn’t IMO capture the nuance of the debate.
Premature decisions about "ethics" and reckless regulation by captive legislators will have an outsized influence on the future, which is why certain players are pushing so hard for both. They appear to have plenty of useful idiots in their corner, judging by how these threads typically play out.
People have no idea what's coming. The personal computer was the biggest deal since electricity. The Internet was the biggest deal since Gutenberg. Large language models are the biggest deal since fucking Aristotle.
Nothing will look the same in 20 years. Absolutely nothing. Whatever you're clinging to now won't be there after the storm, so you'd better learn to swim.
Bullshit
Sony will still do it. It’s about who gets to profit from making humans less valuable as creators, not about whether or not it will happen.
This is about rent seeking by huge corporations, ultimately, since they own huge swaths of rights to creative works and can easily acquire more.
This IS happening. The question is will the benefits to society be democratized to some extent, or will they be concentrated to benefit only the major industry players?
I'm glad you see it that way, and not as some inherent property right. It exists to advance the arts and sciences for humanity, not for the sake of artists. The artists are a means to an end, in this conception!
But if an AI can truly spit out unlimited works that are as good as artists can make, isn't the interest of advancing the arts and sciences accounted for?
And if it can't (what I believe, not from dismissing the models powers but from my understanding of what art is), then do artists have much to fear?
Fair use's factors each weigh for or against a finding of fair use, as opposed to needing to strictly satisfy all four. In particular, what machine learning is likely heavily resting on is that "The more transformative the new work, the less will be the significance of other factors" (Campbell v. Acuff-Rose Music).
For instance, Google Translate was trained on translator's works, and may in part compete with the market for translations, but I'd claim is transformative by nature of adding something new (instant on-demand translation of novel text) and not merely superseding the static works it was trained on.
How to decide "the effect of the use upon the potential market for or value of the copyrighted work" is also a bit of a gray area and one of the questions the US Copyright Office were seeking comments on. Should it be about the impact models have on the market for that general class class of works? Or, the extent to which training on a specific work impacted the market for that specific work compared to if the model was not trained on that work?
It may not be strictly illegal, but in the view of many, including the author of this tweet and myself, it's incredibly unethical to cut artists away from the value they were responsible for, and even more perverse to then use that work to put them out of a job while claiming it's for the good of humanity.
He may want to be an author or musician someday and it is clearly in society's best interest to ensure his mind remains uncorrupted by words and ideas belonging to others.
How much better might the world be had Ford done the same and not endangered the employment of individual manual auto workers with his assembly line, or Gutenberg been more mindful of those poor calligraphers who worked tirelessly for centuries in preserving the books he would otherwise have never been able to print.
The idea that progress is perhaps more important to society's overall health than the preservation of the status quo is such a horrid idea and I mourn the many chiselers who lost their jobs to bronze smithing.
Also, maybe with Neuralink we can finally identify exactly what copyrighted works inform human creativity such that we can properly charge licensing and residuals for such usage.
At least then we might see advancing technology used for something good.
Ignored and shunned? Absolutely.
It's only later on that their focus shifted to where it belonged on the more institutional issues leading to such dire circumstances in the first place.
In the case of the generative AI companies, I think there is a good amount of content that has been consumed that was neither in the public domain or properly licensed.
So what is the legal remedy?
Not even counting watching something that was torrented. Or watched in the wrong country via a VPN. Or a Meme that doesn't pay the original photographer.
The majority of what AI's consume is from publicly available web pages. Putting a "license" requirement on reading those is problematic.
If I watch a movie and it shows an artwork in the background of a scene is that licensed (there are been lawsuits along these lines)? If I hear music being played loudly by someone else is that licensed?
What exactly am I licensed to do with a magazine I find by the side of the road? Or a random webpage I just read?
Lots of edge cases. Copyright Law made a lot more sense when they only people creating and publishing content were a few companies. Not the entire population
So, while some edge cases may technically exist, I wouldn’t call them meaningful in any way whatsoever.
https://www.natlawreview.com/article/who-owns-memes
But if what you have asserted is true then me reading anything that I casually come across on the Internet is okay then it is okay for a web-crawler to do the same and feed that data to an AI.
> Generally, posting and sharing memes online as a form of expression is not actionable and is usually protected under the First Amendment and the doctrines of fair use.
I am saying this is probably the same for humans using the Internet. They don't have an explicit license to read or learn from much of what they read on the Internet.
You are asserting that most of the edge cases are fair use or similar.
The problem is that OP is asserting that an explicit license is required for an AI to learn from content.
My main point is that it is a contradiction to believe that it is legal for humans to read and learn from anything and illegal for machines to do they same. Either it is legal or illegal for both.
Fair use is a defense to all possible claims of copyright infringement, not just the violation of the exclusive right to create derivative works.
I can see the argument to be made, but I personally think they either come from a zero-sum mentality (lack of abundance mentality in this case) or, likely not in this authors' case, or are being called upon as a way to address a different issue (fear, typically)
It's also not technically possible at this point to attribute generated content back to the source training data, so either works will have to be equally compensated (including my meme posted on Reddit) or each piece will need to be individually negotiated with the rights holder. I don't see how these models can exist in these conditions
I think this is the real question. Our culture and ways of life evolved through millennia, and the majority of its bits and pieces were not thought out for dealing with generative AIs. But a bit of thinking tells us that some generative AIs are detrimental to the most cherished aspect of human culture, like music, prose or graphical arts. Perhaps this time we should make an exception and "halt progress" so that we can save our musicians, graphical artists, and novelists. After all, they are our "culture makers", and without them, we are a vessel without a pilot.
In the same breath, I'm going to say that other applications of generative AIs may not deserve the same treatment. For example, if an AI puts out of a work architects, software engineers, medical doctors and so on, and in exchange we get cheaper housing, cheaper software, and eternal youth, then I may wholeheartedly support the AI.
Enforcing copyright on training data won't keep image, music and story generators from being built. It will only stop the open-source ones. OpenAI and Google can afford to license the datasets. It'll just add to their moat!
I can respect wanting to ban image generators in their entirety. I can't respect allowing image generators to be monopolized by the three megacorps who signed a billion-dollar deal with Getty Images.
Humans can also create works that compete with the copyrighted works they've been trained on. Should music schools and conservatoriums start banning all commercial works during lessons? Should students forfeit a % of their lifetime earnings distributed to the copyright holders of everything they have listened to in their lives so far? Granted, nobody will be happy with computers taking over chunks of 'creative' work, but this hard stance will not hold water.
I applaud the courage to stand by personal principles over the soft embrace of a nice job though. Not many people can do it, and we need more of it.
You mean Disney and the other movie studios grabbing copyright for a 100+ years and now wanting to make sure they get their cut from anything even slightly derived from that content?
By the time everyone realised its profiting at society’s cost it was too late. Attempts to put the genie back into the bottle via gdpr etc are futile
Same here. The ship has functionally sailed already. All the LLMs are trained on this.