Stable Audio Open
stability.ai
stability.ai
This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have interesting conversations about the ethics of these things.
In the past I’ve pushed back on people who made the argument that “generative AI intrinsically requires theft from artists”, but the terrible quality of models trained on public domain data made it difficult to make that argument in earnest, even if I knew I was right in the abstract.
What you're left with is a machine that produces "things that strongly resemble the original, that would not have been produced, had you not fed the original into the machine."
The fact that there's no "exact copy inside" the machine seems a lot like splitting hairs; like saying "Well, there's no paper inside the hard drive so the essence of what is copyable in a book can't be in it"
If it output those reviews verbatim, sure I can see the issue, the model is over fitting. But if I tweak the model or filter the output to avoid verbatim excerpts, does an amazon lawyer have a solid footing for a "violation of copyright" lawsuit?
Copyright infringement doesn't require exact copies.
A program that always produces copies is the same as a copy. A program that merely can produce copies categorically is not.
The Library of Babel[1] can produce copyrighted works, and for that matter so can any random number generator, but in almost every normal circumstance will not. The same is true for LLMs and diffusion models. While there are some circumstance that you can produce copies of a work, in natural use that's only for things that will come up thousands of times in its training set -- by and large, famous works in the public domain, or cultural touch-stones so iconic that they're essentially genericized (one main copyrighted example are the officially released promo materials for movies).
Again: Imagine two AI machines, different in one way: One of them has been fed "Article X" and the other hasn't.
You press buttons on the machine(s) in the same way.
The machine that was fed "Article X" spits out something that looks like Article X, and the one that wasn't, doesn't.
The magic inside, I don't think will much matter.
Courts can call experts to testify on matters requiring specialized knowledge or expertise.
- You put thing into the machine
- You press buttons, it makes obvious derivative work
- You don't put thing into the machine, and it can't do that anymore.
There is "something" in there GENERATING COPIES and we see exactly where it came from, even if we can't identify it in the code or whatever.
Honestly, instead of trying to cleanup the output, it's much safer to create a licensed input corpus. People haven't because it's expensive and time consuming. Every time I engage with an AI vendor, my first question is do you indemnify from copyright violations of your output. I was shocked that Google Gemini/Bard only added that this year.
I mean recording a good song is hard. Generating a good song almost impossible. But my gut feeling would've been that recreating a popular song for plausible deniability would be a lot easier.
Same with republishing bestselling books and related media. (I.e. take Lord of the rings and feed it paragraph for paragraph into an LLM that you've prompted to rephrase each to a currently bestselling author.)
While the extremes are obvious, there's a big stretch of gray in the middle. A similar issue occurs in non-AI art, the difference between inspiration and tracing/copying isn't well defined either, but the current method of dealing with that (being on a case-by-case basis and a human judging the difference) clearly cannot scale to the level that many people intend to use these tools.
I am simply not aware of anyone successfully doing this.
Same with the idea of "prompting" and the amount required to generate that copywritten output - again there's the extremes of "The prompt includes copywritten information" to "Vague description".
Arguably some of the same issues exist outside AI, just it's accessibility, scale, and lack of a "Legal Individual" on one side complicates things. For example, if I describe Micky Mouse sufficiently accurately to an artist they reproduce it to the degree it's considered copyright infringement, is it me or the artist that did the infringement? Then what if the artist /had/ seen the previously copywritten artwork, but still produced the same output from that same detailed prompt?
It literally says that within ChatGPT is stored, verbatim, large archives of NY Times articles and that they were able to retrieve them through their API.
Which I agree is problematic, and OpenAI doesn't have the right to disseminate that.
But that doesn't mean OpenAI doesn't have the right to train on it.
Content creators are doing a purposeful slight of hand to confabulate "outputting copyrighted data" with "training on copyrighted data".
It's illegal for me to read an NYT article and recite it from memory onto my blog.
It's not illegal for me to read an NYT article and write my own summary of the article's contents on my blog. This has been true forever and has forever been a staple in new content creation.
I don't think so, I think it's usually argued as two different things.
The "training on copyrighted data" argument is usually that we never licensed this work for this sort of use and it is different enough from previously licensed uses that it should be treated differently.
The "outputting copyrighted data" argument is somewhat like your output is so similar as to constitute a (at least) partial copy.
Another argument is that licensed data is whitewashed by being run through a model. So you could have GPL licensed code that is open source run through a model and then output exactly the same but because it has been outputted by the model it is considered "cleaned" from the GPL restrictions. Clearly this output should still be GPL:ed.
> It's not illegal for me to read an NYT article and write my own summary of the article's contents on my blog. This has been true forever and has forever been a staple in new content creation.
What if I compress the NYT article with gzip? What if I build a LLM model that always replies with the full article within 99% accuracy? Where is the line?
This is not a technical issue, we need to decide on this just like we did with copyright, trademarks, etc. Regardless of what you think this is not a non-issue and we cant use the same rules as we did up until now unless we treat all ML systems as either duplication or humans and neither seems to solve the issues.
I don't think anybody is making that argument. The NY Times claims to have gotten ChatGPT to spit out NY Times articles verbatim but there is considerable doubt about that. Regardless, everyone agrees that a verbatim (or close to) copy is copyright violation, even OpenAI. Every serious model has taken steps to prevent that sort of thing.
I think most people would agree that function is copyrightable if recreated verbatim.
I'm sure if I had a JPEG of some copyrighted raw image it could still be argued that it is the same image. JPEG is imperfect, the result you get is the same every time you open it but it's not the same as the original input data.
ChatGPT would give you the same output every time, and it does if you turn off the "temperature" setting. Introduce a bit of randomness into a JPEG decoder and functionally what's the difference? A slightly different string of tokens for ChatGPT versus a slightly different collection of pixels for a JPEG.
I disagree.
If you can get the model to output an article verbatim, then that article is stored in that model.
Just because it’s not stored in the same format is meaningless. It’s the same content regardless of whether it’s stored as plaintext, compressed text, PDF, png, or weights in a model.
Just because you need an algorithm such as a specialized prompt to retrieve this memorized data, is also irrelevant. Text files need to be interpreted in order to display them meaningfully, as well.
You can't get it to do that, though.[1]
The NYT vs OpenAI case, if anything, shows that even with significant effort trying to get a model to regurgitate specific work, it cannot do it. They found articles it had overfit on due to snippets being reposted elsewhere across the internet, and they could only get it to output those snippets, and not in correct order. The NYT, knowing the correct order, re-arranged them to fit the ordering in the article.
Even doing this, they were only able to get a hundred or so words out of the 15k+ word articles.
No one who knows anything about these models disagrees that overfitting can cause this sort of behavior, but the overwhelming majority of the data in these models is not overfit and they take a lot of care to resolve the issue - overfitting isn't desirable for general purpose model performance even if you don't give a shit about copyright laws at all.
People liken it to compression, like the GP mentioned, and in some ways, it really is. But in the most real sense, even with the incredibly efficient "compression" the models do, there's simply no way for them to actually store all this training data people seem to think is hidden in there, if you just prompt it the right way. The reality is only the tiniest fraction of overfit data can be recovered this way. That doesn't mean that the overfit parts can't be copyright infringing, but that's a very separate argument than the general idea that these are constantly putting out a deluge of copyrighted material.
(None of this goes for toy models with tiny datasets, people intentionally training models to overfit on data, etc. but instead the "big" models like GPT, Claude, Llama, etc.)
1. https://fingfx.thomsonreuters.com/gfx/legaldocs/byvrkxbmgpe/...
> Even doing this, they were only able to get a hundred or so words out of the 15k+ word articles.
OK, that’s less material than I believed, which shows the details matter. But we agree that the overfit material, while limited, is stored in the model.
Of course, this can be (and surely is) mitigated by filtering the output, as long as the product is the output and not the model itself.
I disagree. Granted I'm a layman and not a lawyer so I have no clue how the court feels. But I can certainly make very specialized algorithms to produce whatever output I want from whatever input I want, and that shouldn't let me declare any input as infringing on any rights.
For the reducto ad absurdum example: I demand everyone stops using spaces, using the algorithm 'remove a space and add my copyrighted text' it produces an identical copy of my copyrighted text.
For the less absurd example.. if I took any clean model without your copyrighted text, and brute forced prompts and settings until I produced your text, is your model violating the copyright or is my inputs?
It’s not that clear-cut. It falls into the “Fair use doctrine”The cose 107 of the US copyright law states that the resolutiodepends on>
> (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.
Another thing we need to consider is that the law was redacted with the human mind limitations as a unconcious factor, (i.e not many people would be able to recite War and peace verbatim from memory). This just brings up the fact that copyright law needs a complete re-think.
This does not match my understanding of the information available in the complaint. They might claim they were able to do this, but the complaint itself provides some specific examples that OpenAI and Microsoft discuss in a motion to dismiss... and I think the motion does a very strong job of dismantling that argument based on said examples.
https://fingfx.thomsonreuters.com/gfx/legaldocs/byvrkxbmgpe/...
Would it be possible if Studio Ghibli images had not been used in the training?
otherwise those would just be unknown words, same as asking an artist to do that without any examples.
though I am curious how performance would differ between training on only actual studio Ghibli art, only fan art, or a mix. Maybe the fan art could convey what we expect 'studio Ghibli style' to be even more, whereas actual art from them could have other similarities that that tag conveys.
If the produced work is not a copy, why does it matter if it was generated by a biological brain or by a mechanical one?
> All audio files are licensed under CC0, CC BY, or CC Sampling+.
These explicitly permit derivative works and commercial use.
> Attribution for all audio recordings used to train Stable Audio Open 1.0 can be found in this repository.
So it’s not being glossed over, and licenses are being abided by in good faith imo.
I wish they’d just added a sentence to their press release specifying this, though, since I agree it looks suspect if all you have to go by is that one line.
(Link: https://huggingface.co/stabilityai/stable-audio-open-1.0#dat... )
Not saying PoW is safer (it's not), but less ethical is pretty a bold claim.
To be fair, AI training also has a big carbon footprint, but I feel like the utility provided by AI makes it easier to argue that its usefulness counterbalances its ecological harm.
There's a general consensus that paying someone else to do your dirty work doesn't free you of the moral (or, usually, legal) culpability for the damage done. If you knowingly direct your money towards unethical providers, you are directly increasing the demand for unethical behavior.
(That's assuming that the producers themselves are responsible for the ethics. If a producer is doing its best to convert to clean energy as fast as possible, they may be entirely in the clear but POW would still be unethical. In that scenario POW is placing strain on the limited clean energy supplies, forcing the producer to use more fossil fuels than they'd otherwise need to.)
Solving the past harms means starting from scratch, not giving value to people who previously wasted electricity.
Is this the methane flaring argument, or Peter Thiel’s “windmills in Vermont”?
Do you expect a reply when you start like this?
> Be kind. Don't be snarky. Converse curiously; don't cross-examine.
If you can't follow the guidelines it would probably better to remain quiet.
Assume we have 2 validators in the network; the first one owns 90% of the network, the second one owns 10%. Lets call them Whale and Shrimpy, respectively.
To make the numbers round let's assume total circulating supply of ETH is 100 initially and that the yield resulting from being a validator is 10% per year. After the first year, 10 new ETH will have been minted. Whale would have gotten 9 ETH, and Shrimpy would have gotten 1 ETH. OP is assuming that as 9 is bigger than 1, Whale is getting richer faster than Shrimpy. But, let's look at the final situation globally.
At year 0:
Total ETH circulating supply: 100 ETH
Whale has 90 ETH. Owns 90% of the network.
Shrimpy has 10 ETH. Owns 10% of the network.
At year 1:
Total ETH circulating supply: 110 ETH
Whale has 99 ETH. Owns 90% of the network.
Shrimpy has 11 ETH. Owns 10% of the network.
Whale has exactly the same network ownership after validating for 1 whole year, the network is not centralizing at all! The rich are not getting richer any faster than the poor.
TL;DR: Friends don't let friends skip elementary math classes.
Also, only ~30% tokens are staked. The 30% who chose to stake essentially tax the other 70% in use. Each of the validator do the same amount of work (ok, strictly speaking you get to do more when you have more ETH staked, but being a validator is cheap and does not cost significantly more energy even if you are being selected more frequently because running one proposal is too cheap, that's the whole environmental point, right?) except what they receive is proportioned to how much they stake.
I hate being mean, but sorry, remembering to check one's assumption is a habit I gained after elementary school, so maybe that's too hard for you.
This changes absolutely nothing of the calculation. Furthermore, the change in circulating supply last year was of 0.07%.
> Also, only ~30% tokens are staked.
Correct.
> The 30% who chose to stake essentially tax the other 70% in use.
There is something called opportunity cost. With the existence of liquid staking derivatives the choice to stake or not is one of opportunity cost. Plenty of people may consider the return observed by staking insufficient given the opportunity cost and additional risks. Participating in staking is fully permissionless, stakers are not taxing non-stakers. They are being remunerated for their work.
> Each of the validator do exactly same amount of work (that's the point, right) except what they receive is proportioned to how much they stake.
Incorrect. A staker does proportionate amount of work to its stake. That's why it gets paid more. A staker gets paid for fulfilling its duties as defined in the protocol (attesting, proposing blocks, participating in sync committees). For each of those things there are some rewards and some punishments in case you fail to fulfill them. If a staker has more validators running you simply fulfill more of those duties more often, hence your reward scales linearly with number of validators.
That's just a more polite way to say tax. Being permissionless is cool, but it's still tax in my dict.
> There is something called opportunity cost.
And, who is going to be able to have a larger percentage of their funds staked, a poor or a whale? You need a (mostly) fixed amount of liquidity to use the thing.
> Incorrect. A staker does proportionate amount of work to its stake.
Apologies, I edited my original reply which should answer this.
In short, I don't see anything preventing me to run 10000 validators with 32 ETH each with very similar cost to running just one. It's certainly not linear.
It most certainly is not. They are doing a work for the network and getting remunerated for it. That's not a tax. That's what is commonly referred to as a job. A kid that delivers newspapers over the weekend is not taxing the kid that decides not to. Both make a free decision on what to do with their time and effort given how much it's worth to them. Running a validator takes skill, time, opportunity cost, and you assume certain risks of capital loss. You are getting remunerated for it.
> And, who is going to be able to have a larger percentage of their funds staked, a poor or a whale? You need a (mostly) fixed amount of liquidity to use the thing.
Indeed, the protocol cannot solve wealth inequality. That's an out of protocol issue. It cannot cure cancer either.
> In short, I don't see anything preventing me to run 10000 validators with 32 ETH each with very similar cost to running just one. It's certainly not linear.
There are some fixed costs, indeed. But they are rather negligible. You need a consumer-grade PC (1000 USD) and consumer-grade broadband to solo stake. Or you can use a Liquid Staking Derivative which will have no fixed costs but will have a 10% cut. The curve of APY as a function of stake is very flat. Almost anything else around us has greater barriers of entry or economies of scale.
This is a truth that's fundamental to all types of investing. Advantaged people can set aside millions and not touch it for a year or five or twenty. Disadvantaged people can't invest $20 because there's a good chance they'll need it to buy dinner.
Stocks, bonds, CDs, real estate, it all works like this. You've touched on a fundamental property of wealth.
Wel, but at least in PoW you burn actual money (and in the end, actual resource) proportioned to your profit to keep the network running. In PoS you burn nothing.
And in PoW miners tax 100% of holders.
> what they receive is proportioned to how much they stake
Wealthy miners with state of the art ASICS benefit more than some kid mining at home with an old GPU. Maintenance/cost of mining equipment benefits from economies of scale too.
I hate being mean, but sorry, remembering to check one's assumption is a habit I gained after elementary school, so maybe that's too hard for you.
I'm less happy to pay someone a tax just because they are rich and they did barely anything.
As the operator of a single validator node you can get out of here with that take. I'm using up very significant bandwidth, having to keep a computer running 24/7, updating node and OS software, troubleshoot it after a power or internet outage and at some point I will have to replace the SSD since it is constantly reading and writing and will need replacing after a few years.
Is it a full time job? Absolutely not, but is it free from responsibility? Definitely not. If anything, I could be making more than 3%pa elsewhere if I weren't also in it for ideological reasons.
Why? That's doubly bad for non-mining holders: not only does your share of supply get diluted with newly printed coins, but it also get devalued relative to USD when these coins inevitably get sold to pay expenses
In Ethereum's post-merge world, non-staking holders can keep their share of supply the same (or even have it passively increase) when total supply shrinks. And if the supply does increase by ~0.5-1% and you as a holder aren't okay with that amount of dilution, the barrier of entry to stake profitably and protect your share of supply is much, much lower than the barrier of entry of profitable bitcoin mining.
And the total newly issued coins (which are nominally much lower than pre-merge) have a much lower need to be sold off. If you view issuance as a tax on holders, Ethereum's model wins on all counts
>just because they are rich and they did barely anything
But stakers also "keep the networking running", just like miners under PoW. In both cases, it's gonna be the amount of capital involved that decides how the rewards are proportioned, there's no way around it - these permissionless systems ultimately use the inherent scarcity of economic capital as the anti-sybil mechanism with economic incentives to keep everyone honest. PoS just bypasses the need for burning a huge amount of energy and the embarrassing quantity of single-purpose e-waste to indirectly calculate who has how much at stake. It goes straight to the point: the capital at stake is simply measured in the value of the coin itself instead of external energy/hardware.
On the outside it does kinda look like stakers get rewarded passivly for doing nothing, but there are definitely costs involved, they're just mostly economic instead of physical - think of all the usual risk involved in crypto's volatility, now compound that with slashing risks, illiquidity, opportunity costs – staking yield is like 4-5% atm (and has been down only for quite some time), if you're a billionaire whale you definitely have other investment opportunities available that yield way more than that. I mean just the fact that the net supply growth can go negative shows that even internally in the blockchain itself there can be better things to do with your ETH than stake it; these people aren't burning their ETH on transactions fees for fun, they're actively using their ETH to do stuff that gets them some economic utility.
Good point. It invalidates the "good" part but does not make it doubly bad IMO.
And for ETH, well, I don't think it's about protecting value, it's more about:
> In both cases, it's gonna be the amount of capital involved that decides how the rewards are proportioned, there's no way around it
Yes. The difference is, PoW requires you to BURN resource proportioned to your rewards, while PoS just requires you to HAVE (but not burn) it. This makes a huge difference IMO.
For example, I would consider it more "ethical" (whatever that means) to add a light PoW part (with constant or slowly increasing difficulty, that is chosen to reduce environmental impact) to the ETH PoS protocol as-is: the random-chosen validators have to solve a PoW in addition to make their efforts proportioned to how much they stake, instead of being mostly constant.
I'm not sure how you arrived at this when your initial complaint was the rich get richer with PoS. PoW has much higher costs to participate and after a few years you have more costs when you need to upgrade your mining rigs because they're either burnt out or outcompeted by newer hardware.
The only 'solution' (more a mitigation) to Etherium proof of work's environmental harms is to devalue it.
Unlike your example, this project actually seems to be a net positive for society that wasn't built on top of unnecessary harms.
https://www.youtube.com/@MarekGibney/videos
I find it fascinating that you can put all information needed to recreate a whole complex song into a string like
rough stormy morning car rocks hammering
drum solo roaring dragon downtempo
audiosparx-v2-0 seed 5
This means a whole album of these songs could easily fit into a single TCP/IP packet.If a music genre evolves in which each song is completely defined by its title, maybe it will be called "promptmusic".
I will try the new model with the same prompts and upload the results.
In this case, the whole album could indeed fit into a single TCP/IP packet - because the bulk of information that make up those songs is contained in the model, which weights however many gigabytes it does. The packet carrying your album is meaningless until the recipient also procures the model.
(Tangent: this observation was my first mind. blown. experience when reading GEB over a decade ago.)
Stable Audio Open is an open source text-to-audio model [...]
License: https://huggingface.co/stabilityai/stable-audio-open-1.0/blo... STABILITY AI NON-COMMERCIAL RESEARCH COMMUNITY LICENSE AGREEMENT
Stability are one of the worse offenders for abusing the term "open source" at the moment.No one cares. Not in this space.
Companies are more risk-averse, though, and hobbyists on Reddit don't have the money to do anything serious with this software.
My complaint is about people ignoring the license of open source software made by hobbyists. I disagree with your ignorant "AI fans" generalization.
There are good reasons to use something more aggressive. I'm a big fan of the strict copyleft licenses for this, even if that means companies like Google don't want to that software anymore.
Udio, so far, is undefeated.
And ElevenLabs' music demos were very very impressive, but it's still not released.
https://twitter.com/elevenlabsio/status/1788628175766859891
https://twitter.com/flavioschneide/status/178865450379062893...
https://twitter.com/elevenlabsio/status/1788628178786787822
https://twitter.com/elevenlabsio/status/1788628173367685281
https://twitter.com/elevenlabsio/status/1788628171044053386
https://twitter.com/flavioschneide/status/178867866121578134...
What is ultimately going to be undefeated is training your own model.
https://www.udio.com/songs/ai2uAaBffRGdWdTNNqAbDx
https://www.udio.com/songs/19xQAMG6E1UXG7wNvP7nDW
https://www.udio.com/songs/mPAFYyFgo7Nqjb8ypeFfh9
From scratch?
SHA256: 6049ae92ec8362804cb4cb8a2845be93071439da2daff9997c285f8119d7ea40
I appreciate that the underlying tech is completely different and much more powerful, but it is a pretty strange feeling to find a major AI lab's example sounding so similar to an actual Markov chain MIDI generator I made 14-15 years ago: https://youtu.be/depj8C21YHg?si=74a4DHP14EFCeYrB
(Not that similar, just enough for me to go "huh, what a coincidence").
“The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights.”
This should be standard: commons go in, commons go out.
EDIT: might be cc, non commercial
Text to audio is too limiting. I’d rather input a melody or a drum beat and have the AI compose around it.
Their paper says that they trained it on the Lakh MIDI dataset, and they have a section on potential copyright issues as a result.
Assuming you don't care for legal issues, theoretically you could do: raw signal -> something like Spotify Basic Pitch (outputs MIDI) -> Anticipatory (outputs composition) -> Logic Pro/Ableton/etc + Native Instruments plugin suite for full song
Would be a truly amazing thing for sketching songs! All you need is decent humming/singing/whistling pitch. Hum and generate a bass line, guitar lead, strings, etc. And then sing over it - would make solo musicians able to sketch out a song far easier than transcribing melody to a piano roll.
Yes, this is the reason someone becomes a drummer.
[edit] If not, what free and open ML tools that can be used for restoration / deconvolution / denoising are there?