He’s since founded https://www.fairlytrained.org/
Reference: https://x.com/ednewtonrex
He’s since founded https://www.fairlytrained.org/
Reference: https://x.com/ednewtonrex
Even for rightsholders with tens of millions to hundreds of millions of library items like images or audio snippets, the performance of the encoder or similar feature in text-to-X generative models is too poor on the less than billion tokens of text in the large repositories. This includes Adobe's Firefly.
It is also a misconception that large amounts of similar data, like the kinds that appear in these libraries, is especially useful. Without a powerful text encoder, the net result is that most text-to-X models create things that look or sound very average.
The simplest way to dispel such issues is to publish the architecture of the model.
But anyway, even if it were all true, the only reason we are talking about diffusers, and the only reason we are paying attention to this author's work Fairly Trained, is because of someone training on data that was not expressly licensed.
That’s why it’s important for OpenAI to win the upcoming court cases.
If they lose, they’ll survive. But it will be the end of open model releases.
To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.
> It sounds as if you imply that would be bad. But what if it wasn't?
Entirely possible. The early history of aviation was open source in the sense that many unlicensed people participated, and died. The world is strictly better with licensing requirements in place for that field.
But no one knows. And if history is any guide for software, it seems better to err on freedoms that happen to have some downside rather then clamping down on them. One could imagine a world where BitTorrent was illegal. Or cryptography, or bitcoin.
If you can think of a better example, I’d like to know though. I’ll use it in future discussions. It’s hard to think of good analogies when the tech has new social effects.
There is a real reason why some professions are licenced and others are not.
Your analogy is nonsensical. Not having a better one is irrelevant.
Perhaps a better analogy is movies. At least with acting, you can make your own movies, even if you’re on a shoestring budget. With ML, you quite literally can’t make a useful model. There’s not enough uncopyrighted data to do anything remotely close to commercial models, even in spirit.
You know the word "license" has multiple, dissimilar meanings, right?
kill open source ML -> decrease speed of improvements for some open source ML
It being legal is the only guard against that kind of thing. People will still be angry, but they won’t be so numerous. Right now everyone outside of AI almost universally despises the way AI is trained.
Which means you won’t be able to say that you do open source ML without risking your job. People will be angry enough to try to get you fired for it.
(If that sounds extreme, count yourself lucky that you haven’t tried to assemble any ML datasets and release them. The LAION folks are in the crosshairs for supposedly including CSAM in their dataset, and they’re not even a dataset, just an index.)
I don't agree with this. Most people don't care at all, and at best people would argue about some form of compensation.
Saying "everyone" is unsubstantiated.
I mean... "Everyone was angry at Napster" at the same time "everyone is angry at the MPAA/RIAA"
Or rather, you can, but everyone is free to ignore you. A license without teeth is no license at all. The GPL is only relevant because it’s enforceable in court.
I’m sure some countries will try the licensing route though, so perhaps there you’d be able to make one.
EDIT: I misread you, sorry. You’re saying that if OpenAI loses and license fees become the norm, maybe people will be willing to let their data be used for open source models, and a license could be crafted to that effect.
Probably, yes. But the question is whether there’s enough training data to compete with the big companies that can afford to license much more. I’m doubtful, but it could be worth a try.
The irony of GPL, is that it's validity with respect to users is only now tested in court.
https://www.dlapiper.com/en/insights/publications/2024/01/sf...
And likely proprietary ML as well, hopefully.
(To be clear, I think AI is an absolutely incredible innovation, capable of both good and harm; I also think it's not unreasonable to expect it to play a safer, slower strategy than the Uber "break the rules to grow fast until they catch up to you" playbook.)
I'm all for eliminating copyright. Until that happens, I'm utterly opposed to AI getting a special pass to ignore it while everyone else cannot.
Fair use was intended for things like reviews, commentary, education, remixing, non-commercial use, and many other things; that doesn't make it appropriate for "slurp in the entire Internet and make billions remixing all of it at once". The commercial value of AI should utterly break the four-factor test.
Here's the four-factor test, as applied to AI:
"What is the character of the use?" - Commercial
"What is the nature of the work to be used?" - Anything and everything
"How much of the work will you use?" - All of it
"If this kind of use were widespread, what effect would it have on the market for the original or for permissions?" - Directly competes with the original, killing or devaluing large parts of it
Literally every part of the four-factor test is maximally against this being fair use. (Open Source AI fails three of four factors, and then many users of the resulting AI fail the first factor as well.)
> If they lose, they’ll survive.
That seems like an open question. If they lose these court cases, setting a precedent, then there will be ten thousand more on the heels of those, and it seems questionable whether they'd survive those.
> To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.
You're positioning these as opposed because you're focused on the case of Open Source AI. There are a massive number of Open Source projects whose code is being trained on, producing AIs that launder the copyrights of those projects and ignore their licenses. I don't want Open Source projects serving as the training data for AIs that ignore their license.
That depends on the interpretation of "use", and it would be interesting to read what lawyers think. You learned the language largely from speech and copyrighted works. (All the stories, books, movies, etc. you ever read/heard) When you wrote this comment did you use all of them for that purpose? Is the case of AI different?
To be clear that's a rhetorical question - I don't expect anyone here to actually have a convincing enough argument either way.
Ideas/works/etc literally live rent-free in your head. That doesn't mean they should live rent-free in an AI's neural network.
Changing that should involve actually reducing or eliminating copyright, for everyone, not giving a special pass to AI.
Human brain most definitely is not exempt. If you read Lord of the Rings and then write down a new book, with the same characters and same story line - that's plain copying(lookup the etymology of the verb to copy). If you look at a painting and paint a very similar painting - that's still copying.
Human brains are the reason we have copyright. Your recital of passages from any copyrighted book would violate the copyright, if not for fair use doctrine. And it has nothing to do with whether you do it yourself, or have a TTS engine produce the sound.
I'm saying that AI does not and should not automatically get the exception that a human brain does.
There’s also the moral question. Should creators have the right to prevent their bits from being copied at all? Fundamentally, people are upset that their work is being used. But "used" in this case means "copied, then transformed." There’s precedent for such copying and transformation. Fair use is only one example. You’re allowed to buy someone’s book and tear it up; that copy is yours. You can also download an image and turn it into a meme. That’s something that isn’t banned either. The question hinges on whether ML is quantitatively different, not qualitatively different. Scale matters, and it’s a difference of opinion whether the scale in this case is enough to justify banning people from training on art and source code. The courts’ opinion will have the final say.
The thing is, I basically agree with you in terms of what you want to happen. Unfortunately the most likely outcome is a world where no one except billion dollar corporations can afford to pay the fees to create useful ML models. Are you sure it’s a good outcome? The chance that OpenAI will die from lawsuits seems close to nil. Open source AI, on the other hand, will be the first on the chopping block.
really it seems more like someone was afraid of angering Nintendo who is a corporate adversary one does not like to fight and thus it has a bunch of blocks to keep from generating anything that offends Nintendo, that does not really translate to quickly and easily stopping and blocking offending generations across every copyrighted work in the world.
What I don't understand (as a European with little knowledge of court decisions on fair use): with the same reasoning you might make software piracy a case of 'fair use', no? You take stuff someone else wrote - without their consent - and use it to create something new. The output (e.g. the artwork you create with Photoshop) is definitely not copyrighted by the manufacturer of the software. But in the case of software piracy, it is not about the output. With software, it seems clear that the act of taking something you do not have the rights for and using it for personal (financial) gain is not covered by fair use.
Why can OpenAI steal copyrighted content to create transformative works but I cannot steal Photoshop to create transformative works? What am I missing?
That's not a good example. Making a copy of a record you own(as an example ripping a audio CD to MP3) is absolutely fair use. Giving your video game to your neighbor to play - that's also fair use.
Fair use is limited when it comes to transformative/derivative work. Similar laws are in place all over the world, just in US some of those come from case law.
> With software, it seems clear that the act of taking something you do not have the rights for and using it for personal (financial) gain is not covered by fair use.
> Why can OpenAI steal copyrighted content to create transformative works but I cannot steal Photoshop to create transformative works?
That's not a good analogy. The argument, that is not settled yet, is that a model doesn't contain enough copyrightable material to produce an infringing output.
Take your software example - you legally acquire Civ6, you play Civ6, you learn the concepts and the visuals of Civ6... then you take that knowledge and create a game that is similar to Civ6. If you're a copyright maximalist - then you would say that creating any games that mimic Civ6 by people who have played Civ6 is copyright infringement. Legally there are definitely lower limits to copyright - like no one owns the copyright to the phrase "Once upon a time", but there may be a copyright on "In a galaxy far far away".
If Photoshop was hosted online by Adobe, you would be free to do so. It's copyrighted, but you'd have an implied license to use it by the fact it's being made available to you to download. Same reason search engines can save and present cached snapshots of a website (Field v. Google).
In other situations (e.g: downloading from an unofficial source) you're right that private copying is (in the US) still prima facie copyright infringement. However, when considering a fair use defense, courts do take the distinction into strong consideration: "verbatim intermediate copying has consistently been upheld as fair use if the copy is ‘not reveal[ed] . . . to the public.’" (Authors Guild v. Google)
If you were using Photoshop in some transformative way that gives it new purpose (e.g: documenting the evolution of software UIs, rather than just making a photo with it as designed) then you may* be able to get away with downloading it from unofficial sources via a fair use defense.
*: (this is not legal advice)
https://www.copyright.gov/title17/92chap1.html#107
(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
(2) the nature of the copyrighted work;
(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
(4) the effect of the use upon the potential market for or value of the copyrighted work.
There's more detailed discussion here: https://copyright.columbia.edu/basics/fair-use.html
If I paint a picture inspired by Starry Night(Van Gogh) - does that inherently infringe on the original? I looked at that painting, learned the characteristics, looked at other similar paintings and painted my own. I basically trained my brain. (and I mean the copyright, not the individual physical painting)
And I mean cases where I am not intentionally trying to recreate the original, but doing a derivative(aka inspired) work.
Because it's already settled that recreating the original from memory will infringe on copyright.
Your first factor seems to not at all be like that which Stanford has in its guidelines[1], which they call the transformative factor:
In a 1994 case, the Supreme Court emphasized this first factor as being an important indicator of fair use. At issue is whether the material has been used to help create something new or merely copied verbatim into another work.
LLMs mostly create something new, but sometimes seems to be able to regurgitate passages verbatim, so I can see arguments for and against, but to my untrained eyes doesn't seem as clear cut.
[1]: https://fairuse.stanford.edu/overview/fair-use/four-factors/
In the broadest sense, generative AI helps achieve the same goals that copyleft licences aim for. A future where software isn't locked away in proprietary blobs and users are empowered to create, combine and modify software that they use.
Copyleft uses IP law against itself to push people to share their work. Generative AI aims to assist in writing (or generating) code and make sharing less neccesary.
I argue that if you are a strong believer in the ultimate goals of copyleft licences you should also be supporting the legality of training on open source code.
If an artist approached a software developer, created a painting of them using their Mac, and said "There, I've done your job for you" you'd think they were an idiot.
This is the same from the other side. The inability to understand why that's a realistic analogy does not change the fact that it is.
What a curious type of theft where the author keeps their art and I get different art.
This is why it is important whether you consider that infringement occurs upon ingestion or output. If it only matters for outputs, then artists have a problem, since copyright doesn't protect styles at all, see for example the entire fashion industry.
There is a saving grace though: Artists can make a case that the association of their distinctive style with their name is at least potentially a violation of trademark or trade dress, especially if that association is being used to promote the outputs to the public. This is a fairly clear case of commercial substitution in the market for creating new works in that artist's style and creating confusion concerning the origin of the resulting work.
Note that the market for creating new works in a particular artist's distinctive and named style kind of goes away upon the artist's passing. What remains is the trademark issue of whether a particular work was actually created by the artist or not, which existing trademark law is well suited to policing, as long as the trademark is defended, even past the expiration of the copyright.
Meanwhile, trademark (and copyright) also apply to the subjects of works, like Nintendo's Mario or Disney's Mickey Mouse or Marvel's Iron Man. But we don't really want models to simply be forbidden from producing them as outputs, or they become useless as tools for the purpose of parody and satire, not to mention the ability to create non-commercial fan art. The potential liability for violating these trademarks by publishing works featuring those characters rests with the users rather than the tools, though, and again existing law is fairly well suited to policing the market. Similarly, celebrities' right of publicity probably shouldn't prevent models from learning what they look like or from making images that include their likeness when prompted with their name, but users better be prepared to justify publishing those results if sued.
You can also make the (technical) argument that if you just ask for an image of Wonder Woman, and you get an image that looks like Gal Gadot as Wonder Woman, that the model is overfitting. That's also the issue with the recent spate of coverage of Midjourney producing near-verbatim screenshots from movies.
It might be appropriate though to regulate commercial generative AI services to the extent of requiring them to warn users of all the potential copyright/trademark/etc. violations, if they ask for images of Taylor Swift as Elsa, or Princess Peach, or Wonder Woman, for example.
This is contrary to the goals of the Free Software movement - and also why Free Software people were the first to complain about all the copying going on. One of the things Generative AI is really good at is plagiarism - i.e. taking someone else's work and "rewriting it" in different words. If that's fair use, then copyleft is functionally useless.
It's important to keep in mind the difference between violating the letter of the law and opposing the business interests of the people who wrote the law. Copyleft and share-alike clauses have the intention of getting in the way of copyright as an institution, but it also relies on copyright to work, which is why the clauses have power even though they violate the spirit of copyright. Generative AI might violate the letter of the law, but it's very much in the spirit of what the law wants.
[0] Cory Doctorow: "Intellectual property is any law that allows you to dictate the conduct of your competitors"
Creative Commons has been fairly pro-AI -- they have been quite balanced, actually, but they do say that opt-in is not acceptable, it should be opt-out at most. EFF is fairly pro AI too -- at least, against using copyright to legislate against it.
You shouldn't discount progress in the open model ecosystem. You can sort of pirate ChatGPT by fine tuning on its responses, there's GPU sharing initiatives like Stable Horde, there's TabbyML which works very well nowadays, and Stable Diffusion is still the most advanced way of generating images. There's very much of an anti-IP spirit going on there, which is a good thing -- it's what copyleft is there for in sprit, isn't it?
The Software Freedom Conservancy has been complaining about GitHub Copilot since 2022[0]. They specifically cite Copilot's use of training data in ways that violate the copyleft and attribution requirements of various FOSS licenses. Hector Martin (the guy porting Linux to MacBooks) also agrees with this. It's also important to note that the first AI training lawsuit was specifically to enforce GPL copyleft[1].
The EFF's argument has come across to me less like "AI is cool and good" and more like "copyright doesn't do a good job of protecting artists against AI taking their jobs". Cory Doctorow's also taken a similar position, arguing that unions are better at protecting against AI than copyright is. e.g. WGA being able to get contractual provisions preventing workers from being replaced with AI.
This is a different vein of opposition to AI from what we saw the following year in 2023 with artists and writers, though. Even then, those artists and writers aren't suddenly massively pro-copyright[2] and more consider it a means to fatally wound AI companies[3]. In contrast, big businesses that own shittons of copyright have been oddly quiet about AI. Sure, you have Getty Images and The New York Times suing Stability and OpenAI, but where's, say, Disney or Nintendo's litigation? These models can draw shittons of unlicensed fanart[4], and nobody cares. Wizards and Wacom made big statements against AI art, but then immediately got caught using it anyway, because stock image sites are absolutely flooded with it.
My personal opinion is that generative AI creates enough issues that we can't group them down into neat "pro-copyright" vs. "anti-copyright" arguments. People who share their work for free online are complaining about it while people who expect you to pay money for their work are oddly ambivalent. AI is orthogonal to copyright.
I will give you that the open model community is doing cool shit with their stolen loot. However, that's still something large corporations can benefit from (e.g. Facebook and LLaMA).
[0] https://sfconservancy.org/GiveUpGitHub/
[1] https://en.wikipedia.org/wiki/GitHub_Copilot#Licensing_contr...
[2] Which, for the record, many of them break.
[3] Their actual argument against AI is based on moral grounds, not legal ones. I don't think any artist is going to accept licensing payments for training data, they just want the models deleted off the Internet, full stop.
[4] OpenAI tried to ban asking for fanart, but if you ask for something vaguely related (e.g. "red videogame plumber" or "70s sci-fi robot") you'll get fanart every time.
"many other things" has included, for example, Google Books scanning millions of in-copyright books, storing internally them in full, and making snippets available.
The basis for copyright itself is to "promote the progress of science and useful arts". For that reason a key consideration of fair use, which you've skipped entirely, is the transformative nature of the new work. As in Campbell v. Acuff-Rose Music: "The more transformative the new work, the less will be the significance of other factors", defined as "whether the new work merely 'supersede[s] the objects' of the original creation [...] or instead adds something new".
> "How much of the work will you use?" - All of it
For the substantiality factor, courts make the distinction between intermediate copying and what is ultimately made available to the public. As in Sega v. Accolade: "Accolade, a commercial competitor of Sega, engaged in wholesale copying of Sega's copyrighted code as a preliminary step in the development of a competing product" yet "where the ultimate (as opposed to direct) use is as limited as it was here, the factor is of very little weight". Or as in Authors Guild v. Google: “verbatim intermediate copying has consistently been upheld as fair use if the copy is ‘not reveal[ed] . . . to the public.’”
The factor also takes into account whether the copying was necessary for the purpose. As in Kelly v. Arriba Soft: "If the secondary user only copies as much as is necessary for his or her intended use, then this factor will not weigh against him or her"
While there are still cases of overfitting resulting in generated outputs overly similar to training data, I think it's more favorable to AI than simply "it trained on everything, so this factor is maximally against fair use".
> Directly competes with the original, killing or devaluing large parts of it
The factor is specifically the effect of the use upon the work - not the extent to which your work would be devalued even if it had not been trained on your work.
Substantiality of code does not apply to substantiality of style. What's being copied is look and feel, which is very much protected by copyright.
The copying clearly is necessary for the purpose. No copying, no model. The fact that the copying is then compressed after ingestion doesn't change the fact that it's necessary for the modelling process.
Last point - see first point.
IANAL, but if I was a lawyer I'd be referring back to look and feel cases. It's the essence of an artist's look and feel that's being duplicated and used for commercial gain without a license.
That's true whether it's one artist - which it can be, with added training - or thousands.
Essentially what MJ etc do is curate a library of looks and feels, and charge money for access.
It's a little more subtle than copying fixed objects, but the principle remains the same - original work is being copied and resold.
The question for transformative nature is whether it merely supersedes or instead adds something new. E.G: Google translate was trained on books/documents translated by human translators and may in part displace that need, but adds new value in on-demand translation of arbitrary text - which the static works it was trained on did not provide.
> Substantiality of code does not apply to substantiality of style.
I'm not certain what you're saying here.
> The copying clearly is necessary for the purpose. No copying, no model.
Which, for the substantiality factor, works in favor of the model developers.
> It's the essence of an artist's look and feel that's being duplicated and used for commercial gain without a license.
Copyright protects works fixed in a tangible medium, not ideas in someone's head. It would protect a work's look/appearance (which can be an issue for AI when overfitting causes outputs that are substantially similar to a protected work), but not style or "an artist's look and feel".
If that were the case, no one would be able to paint any cubist paintings. (Picasso estate would own the copyright, to this day)
It's not that clear cut, there are a lot of nuances.
That succeeds on a different part of the four-factor test, the degree to which it competes with / affects the market for the original.
Google Books is not automatically producing new books derived from their copies that compete with the original books.
It satisfied multiple parts of the four-factor test. It was found satisfy the first factor due to being "highly transformative", the second factor was considered not dispositive is isolation and favoring Google when combined with its transformative purpose, and it satisfied the third factor as the usage was "necessary to achieve that purpose" - with the court making the distinction between what was copied (lots) and what is revealed to the public (limited snippets).
As you had all factors as "maximally against" fair use, do you believe that AI is significantly less transformative than Google Books? I'd say even in cases where the output is the same format as the content it was trained on, like Google Translate, it's still generally highly transformative.
> the degree to which it competes with
Specifically, to be pedantic, it's the effect of the use/copying of the original copyrighted work.
Genuinely curious: for anyone who thinks AI obviously violates copyright, how do you resolve this? E.g. do you think the violation happens during training or inference? And is it the trained model, or the model output, that you think should be considered a derived work?
Just like a translation of a book is a derived works of the original. Or a binary compiled output is a derived works of some source code.
> In copyright law, a derivative work is an expressive creation that includes major copyrightable elements of ... the underlying work
A trained model fails that on two counts, doesn't it? Both the "includes" part, and the fact that a model is itself not an expressive work of authorship.
There's nothing creative about the act of a compiler, it is automatic, just like the training run of an LLM.
And no part of the original source code is in the binary output.
And yet, binaries are a derived work from the source code that went into them.
So something is up! I am not a lawyer though.
It's not about whether the binary includes the raw text of the source, but whether it copies the expressive content. Anything expressive (i.e. copyrightable) in a compiled binary must have come from the sourcecode, so that's what makes it a derived work.
But the same isn't true of LLMs, which are more like "data about their inputs", than "a transformed version of their inputs".
Translation of a book is non-transformative and retains the original author's artistic expression.
As a counter example - if you write an essay about Picasso's Guernica painting, it is derivative according to our colloquial use of the term, but legally it's an original work.
If the US had tighter regulations, China or someone else will take over the market. If AI is genuinely transformative for productivity, then the US would just fall behind, sooner or later.
Like, we see this line everywhere now, and it simply doesnt make sense. At some point you just have to believe something, be principled. Treating the entire world as this zero sum deadlock of "progress" does nothing but prevent one from actually being critical about anything.
This would-be Oppenheimer cosplay is growing really old in these discussions.
The alternative is that any models widely trained on copyrighted work are uncopyrightable and must be disclosed, along with their data sources. In essence this is forcing all such models to be open. This is the only equitable outcome. Any use of the model to create works has the same copyright issues as existing work creation, ie if substantially replicates an existing work it must be licenced.
All outcomes suck. The trick is to find the outcome that sucks the least for the majority of people. Maybe the needs of copyright holders will outweigh the needs of open source, but it’s basically guaranteed that open source ML will die if your first paragraph comes true.
We might apply that as a $5000 or so surcharge on AI accelerators capable of running the models, such as the 4090.
Absolutely true. That's the end game and we should be working toward influencing that. It's within our power.
> I’ve spoken with a few lawyers who believe OpenAI is on solid legal footing
No one knows anything, this is too novel, and even if OpenAI gets some fair use ruling, it will be inequitable and legislation is inevitable. OpenAI is between a rock and a hard place here. If you read the basis for fair use and give each aspect serious consideration, as a judge should do, I can't see it passing fair use muster. It's not a case of simply reproducing work, which in unclear here, it's the negative effect on copyright holders, and that effect is undeniable.
> All outcomes suck.
I don't think so. It's possible to fashion something equitable, but people other than the corporations have to get involved.
Whether or not that’s equitable is in the eye of the beholder. Copyright is an artificial construct, not a natural law. There is nothing that says we must have it, or we must have it in its current form, and I would argue the current system of copyright has been largely harmful to creativity for a long time now. One of the most damning statements I’ve read in this thread about the current copyright system is how there’s simply not enough unlicensed content to train models on. That is the bed that the copyright-holding corporations have made for themselves by lobbying to extend copyright to a century, and it all but assured the current situation.
No I'm saying that's what they law should be, because models can be built and used without anyone knowing. If it's illegal not to disclose them you can punish people.
Copyright is something that protects the little guy as much as big corps. But the former has more to lose as a group in the world of AI models, and they will lose something here no matter what happens.
I'd love to hear that argument.
How has the current system of copyright been harmful to creativity?
It is possible to strangle OpenAI without strangling AI: pmarca is anti-OpenAI in print, but you can bet your butt he hopes to invest in whatever replaces it, and he’s got access to information that like, 10 people do.
A useful example would be the Napster Wars: the music industry had been rent seeking (taking the fucking piss really) for decades and technology destroyed the free ride one way or another. The public (led by the technical/hacker/maker public) quickly showed that short of disconnecting the internet, we were going to listen to the 2 good songs without buying the 8 shitty ones. The technical public doesn’t flex its muscles in a unified way very often, but when it does, it dictates what is and isn’t on the menu.
The public wants AI, badly. They want it aligned by them within the constraints of the law (which is what “aligned” should mean to begin with).
The public is getting what it wants on this: you can bet the rent. Whether or not OpenAI gets on board or gets run the fuck over is up to them.
“You in the market for a Tower Records franchise Eduardo?”
This is another one of those “well if you treat the people fairly it causes problems” sort of arguments. And: Sorry. If you want to do this you have to figure out how to do it ethically.
There are all sorts of situations where research would go much faster if we behaved unethically or illegally. Medicine, for example. Or shooting people in rockets to Mars. But we can’t live in a society where we harm people in the name of progress.
Everyone in AI is super smart — I’m sure they can chin-scratch and figure out a way to make progress while respecting the people whose work they need to power these tools. Those incapable of this are either lazy, predatory, or not that smart.
As an ML researcher, no, there’s basically no way to make progress without the data. Not in comparison with billion dollar corporations that can throw money at the licensing problem. Synthetic data is still a pipe dream, and arguably still a copyright violation according to you, since traditional models generate such data.
To believe that this problem will just go away or that we can find some way around it is to close one’s eyes and shout "la la la, not listening." If you want to kill open source AI, that’s fine, but do it with eyes open.
Also, beware of originalist interpretations of the Constitution. I believe there’s been about 250 years of law clarifying how copyright works, and, not to beat a dead horse, I don’t think it carves out a special exception for open source projects.
I don't think this is true. There's a huge amount of public domain works, as well as stuff licensed under permissive copyleft licenses, that can be used.
But, even if it did kill off open-source ML, it would still be necessary, because it's morally wrong to train ML models on copyrighted content without compensating the copyright owners (on their terms).
As a content creator, I explicitly do not want or consent to any of my creative works being used to train ML models without having a licensing agreement through which I am financially compensated.
Morals are different from the law, but you seek a legal remedy, and those aren’t going well.
> failed to provide evidence supporting any of their claims except for direct copyright infringement
(emphasis mine)
Where are the courts saying that models can be trained on copyrighted content? (I believe that it's possible but unless I'm missing something I don't see it in that Ars article)
Thanks for putting this into words. I'm of the same opinion and this is the best articulation I have so far.
I think that also helps in understanding his departure, since he's a founder with a music background.
Person B gets hired to observe Person A working, check email, and be the audio output buffer for Jira.
Person B says "I built this."
That's dishonesty no matter what the titles are or how important the emails were.
A captain may steer the ship, but they're not the one actually creating and maintaining the means by which it moves.
The crux of the debate appears to rest on the fact that context matters. If an engineer says to another engineer "I built the spam detection system", it is understood that they mean they either wrote the code or had some direct part in producing it. If an executive says to another executive "I made the Mac", neither is interpreting that as them literally building the thing. They know they are in leadership, the meaning is assumed to be "as a leader".
And yet virtually everyone will go along with a statement like "The captain sailed the ship across the ocean" or "Captain Kirk charted the Gamma Quadrant" or whatever, so I'm not sure how this serves as an objection to the original phrasing.
And to be clear, I’m not sure Ed would call himself that. Those are my words, not his.
Agreed, I wouldn't say I was hired to build Stable Audio. Crazy talented team of research engineers / software engineers / designers did the building.
Also wanted to clarify that I didn't quit due to concerns around the training data used for Stable Audio. I was proud of the approach we took to training data - a rev share with rights holders. I quit because of the prevailing view on training data at the wider company, as documented in its public response to the copyright office, where it argues that training on people's work without consent is fair use.
I have spent more blood, tears and money on art than most of you would find even remotely bearable.
I not only consider my songs to be fair use for training a model but I would also honored if my works were included and influenced further musicians in a way that my records probably never will.
The best songwriters I know have other careers and keep on going otherwise. If you actually care about musicians you should make it a habit to go see local live music!
Also congrats on the new company!
Or did he needed that as it i part of the business model of his certfications?
Ed still likes Stability, especially as we fully trained stable audio on rights licensed data (bit different in audio to other media types), offer opt out of datasets etc.