Open-sourcing AudioCraft: Generative AI for audio
ai.meta.com
ai.meta.com
Meta is really clearly trying to differentiate themselves from OpenAI here. Open source + driving home "we don't use data we haven't paid for / don't own".
Furthermore, MusicGen's weights are licensed CC-BY-NC, which is effectively a nonlicense as there is no noncommercial use you could make of an art generator[1]. This is not only a 'weights-available' license, but it's significantly more restrictive than the morality clause bearing OpenRAIL license that Stability likes to use[2].
[0] https://github.com/facebookresearch/llama/blob/main/MODEL_CA...
[1] https://github.com/facebookresearch/audiocraft/blob/main/LIC...
[2] These are also very much Not Open Source™ but the morality clauses in OpenRAIL are at least non-onerous enough to collaborate over.
How do you figure? Have you never just...made stuff to make stuff?
Obviously, you can't host a commercial art generation service with a noncommercial-use license, and (insofar as art produced by a generator is a derivative work of the model weights, which is a controversial and untested legal theory) you can’t make commercial art with a noncommercial license, but not all art is commercial.
You're probably thinking of "not charging a fee to use", which is a subset of all the ways you can monetize a creative work. You can still make money off of AudioCraft by just hosting it with banner ads next to the output. Even a "no monetization" clause[0] would be less onerous than "noncommercial use only", because it'd at least be legal to use AudioCraft for things like background music in offices.
[0] Which already precludes the use of AudioCraft music on YouTube since you can't do unmonetized uploads anymore
The definition of “NonCommercial”, the oddly capitalized term of art in the license, is not a matter of general law, it is a matter of the license, which defines it as “not primarily intended for or directed towards commercial advantage or monetary compensation. For purposes of this Public License, the exchange of the Licensed Material for other material subject to Copyright and Similar Rights by digital file-sharing or similar means is NonCommercial provided there is no payment of monetary compensation in connection with the exchange.”
> Even if you don’t intend to make money the law still considers the work itself to be commercial.
Even if you do make money, if the use is “not primarily intended” for that purpose, it is "NonCommercial" in the terms of the license.
> That’s why CC-BY-NC has to have a special “filesharing is non-commercial” statement in it, because people have made successful legal arguments that it is.
It has the filesharing term in it because it permits that particular exchange-of-value as a primary purpose.
> Even a “no monetization” clause would be less onerous than "noncommercial use only"
How would a clause that prohibits monetization entirely be less onerous than one which prohibits it only as the primary intent of use?
> it’d at least be legal to use AudioCraft for things like background music in offices.
It is legal to use it for that purpose (in a for-profit enterprise, I suppose, one might make an argument that any activity was ultimately primarily directed at “commercial advantage”, but in a government or many nonprofit environments, that wouldn’t be the case.)
I realize, this isn't legal advice, YMMV, etc.
A resort, probably not, ambiance is, at least arguably, a marketable commercial advantage; a private club in the “mutual benefit organization” sense (rather than a “business selling memberships”, which is just like a resort), probably, because their interest, even indirectly, isn’t making money.
- If I use AudioCraft to post freely-downloadable tracks on my SoundCloud, I still get the benefit of having a large audio catalog in my name, even if I'm not selling the individual tracks. I could later compose tracks on my own and ride off the exposure I got from posting "noncommercially".
- If I run AudioCraft as a background music generator in my store, I save money by not having to license music for public performance.
- If I host AudioCraft on a website and put ads on it, I'm making money by making the work available, even though I'm not charging a fee for entry.
I suspect that a lot of people reading this are going to have different arguments for each. My point is that if you don't think that all of these situations are equally infringing of CC-BY-NC, then you need to explain why some are commercial and some are not. Keep in mind that every exception you make can be easily exploited to strip the NC clause off of the license.
If you're angry at the logic on display here, keep in mind that this is how judges will construe the license, and probably also how Facebook will if you find a way to make any use of their AI. The only thing that stops them from rugpulling you later is explicit guidance in CC-BY-NC. Unfortunately, the only such guidance is that they don't consider P2P filesharing to be a commercial use.
So, absent any other clarifications from Facebook, all you can do without risking a lawsuit is share the weights on BitTorrent.
EDIT: And yes, I have made stuff just to make stuff. I license all of that under copyleft licenses because they express the underlying idea of 'noncommercial' better than actual noncommercial clauses do.
Do you think that non commercial use simply doesn't exist or something?
Because non commercial use isn't some crazy concept. It is a well established one, that doesnt disclude literally everything.
Also, you are ignoring the idea that Facebook will almost certainly not sue anyone for using this for any reason, except possibly Google or Apple.
So if you aren't literally one of those companies you could probably just use it anyway, ignore the license completely, and have zero risk of being sued.
Whatever happened to esr? Did he just get too paranoid and clam up?
I presume you mean in USA, because in UK you don't have a general private right to copy. Our "Fair Dealing" is super restrictive compared to Fair Use.
I like that it makes software like iTunes contributory infringers for enabling mass copyright infringement.
As for "Facebook won't sue"? Sure, except we don't have to worry about just Facebook. We have to worry about anyone with a derivative model. There's an entire industry of copyleft trolls[0] that could construct copyright traps with them.
Individuals can practically ignore NC mainly because individuals can practically ignore most copyright enforcement. This is for the same reason why you can drive 55 in a 30mph zone and not get a citation. It's not that speeding is now suddenly legal, it's that nobody wants to enforce speed limits - but you can still get nailed. The moment you have to worry about NC, there is no practical way for you to fit within its limits.
[0] https://www.techdirt.com/2021/12/20/beware-copyleft-trolls/
No, for “NonCommercial”, what actually matters is the explicit definition in the license.
Noncommercial licenses are taken up in "GREAT MINDS v. FEDEX OFFICE AND PRINT SERVICES, INC 886 F.3d 91 (2nd Cir. 2018). Thé court explains they are enforceable and are basically just a category of contract. So, as long as the contract is clear, it’s probably enforceable.
You're right: those are all equally infringing CC-BY-NC. I don't see a problem.
> this is how judges will construe the license
What “NonCommercial” means in the license is explictly defined in the license, and if you think either those examples, or more to the point, every possible use ever so as to render ‘NonCommercial’ into ‘no use’ as you have claimed, you need to make that argument, based on the definition in the license, not some concept of what might be construed as commercial use by general legal principles if the license used the term without its own explicit definition.
[1] https://opensource.stackexchange.com/questions/12070/allowed...
I suppose your point would stand if the software were a quine?
I'm sorry, what?
r/stablediffusion gives you a hundred examples daily of people just having fun and not thinking of monetizing their generations
Isn't Meta settling lawsuits for this right now? In addition to violating user privacy (another lawsuit)...
Meta is attempting to destroy competition; that's it. Similar to how they paid a fortune to lobby against Tiktok for the exact reasons Meta is under active investigation (again). The irony.
In the new world that Meta sees, of VR/AR and AI, Meta is in a position already were people don't want them to have much power in this world, because they don't trust them over privacy etc, meta is trying to pivot to become more trustworthy so they make genuine moves in this space.
But their internal research stays internal. Sometimes, they put out "papers" which are glorified advertisements, often going as far as hiding the model architecture just to keep their competitive advantage.
https://www.americanbar.org/groups/science_technology/public...
We need to do better than to repeat these claims uncritically. The weight licenses are not "open source" by any useful definition, and we should not give Meta kudos for their misleading PR (especially considering that they almost surely ignored any copyright when training these things - rules for thee, but not for me).
"Not as closed as OpenAI" is accurate, but also damning with faint praise.
In reality, you can't, as they licensed the weights for noncommercial use only: https://github.com/facebookresearch/audiocraft#license
If you want to build a company, perhaps you should do what everyone in the industry has done for millennia, copy the movements performed and optimize them while doing so.
about: pytorch @ fb.More popular opinion is OSI: https://en.wikipedia.org/wiki/Open_Source_Initiative
They were founded by the persons who (claimed to have) invented the term in order to steward it. It's the same definition as the FSF.
Before, people most often used "free software" as defined by the free software movement, but some disliked this term because it's confusing (most think "free" means no money) and perceived to be anti-commercial.
The term "open source software" was chosen and given a precise definition.
It's dishonest, then, for people to use the term "open source software" with a different interpretation when it was specifically chosen to avoid confusion.
I disagree. You're saying that they "invented" the term, but it's a very generic term. The source is open, it's open source. I bet people were using the term before they claim they invented it.
In that context, it is very fine to use a different definition and in fact here's my definition, and I guess most people (maybe not on HN) share it: if the source is visible by the general public, it's open source.
For what you mean, I use "FLOSS".
The article says this:
> Our audio research framework and training code is released under the MIT license to enable the broader community to reproduce and build on top of our work
IMO, while I'd rather have one part permissively licensed than nothing at all... it stinks that companies sponsoring researchers get an un-nuanced level of street cred for "open sourcing" something that they know nobody will ever be able to reproduce because their data set and/or their compute grid's optimizations are proprietary.
As it stands, I'm not at all sure that the outputs of this model can be used for commercial videos.
Maybe it's not a big deal to 'lose' the past, maybe landfills will be mined for authentic content.
Fortunately, it's pretty simple in real life. We have certain publications and sources we trust, whether they're the NYT or a respected industry blog. We know they take accurate reporting seriously, fire journalists who are caught fabricating things, etc.
If we see a clip on YouTube from the BBC, we can trust it's almost certainly legit. If it's some crazy claim from a rando and you care whether it's real, it's easy to look it up to see if anyone credible has confirmed it.
So no, no worry at all about the past being erased.
This goes all the way back to yellow news with newspapers: https://en.wikipedia.org/wiki/Yellow_journalism
Imagine facebook decides to subtly change every public post and comment to show some particular person or cause in a better light.
We have tons of credible archived sources owned by different institutions. And these sources are successful in large part due to their credibility and trustworthiness.
It's just not economically rational for any of them to start "altering the past", and if they did, they'd be caught basically immediately and their reputation would be ruined.
This isn't an ML/tooling question, it's a question of humans and reputation and economic incentives.
Second, people take screenshots of Facebook posts all the time. They're everywhere. If you suddenly have a ton of people with timestamped screenshots from their phones that show Facebook has changed content, that's exactly the kind of story journalists will pounce on and verify.
The idea that Facebook could or would engage in widespread manipulation of past content and not get caught is just not realistic.
Maybe it is improbable, but there now is the technical possibility which was not there before.
It is valuable to explore that possibility and maybe even work to prevent such a use.
I would be interested in a ledger of cryptographically signed records of important public information such as newspapers, government communication and intellectual discourse.
Your argument that large social media will behave rationally is not backed up by reality. Consider Musk and Twitter.
Detection doesn't really matter, because people are too lazy to validate the facts, and reporters are not interested in reporting them. AI is simply another tool to manipulate people, like Wikipedia, Reddit.com, Twitter, or any other BS psuedo-authority. Think someone will actually crack open a book to prove the AI wrong? Not a chance.
You really think that if the NYT started altering its past stories, other publications would just... ignore it?
It would be a front-page scandal that the WaPo would be delighted to report on. As well as a hundred other news publications.
Thankfully.
If you can't alter world news headlines, you can still alter the tone of the article. If you can't alter front page news, you still can alter the remaining 95% of news.
Influencing public opinion is more subtle than the one important headline per day.
You are also ignoring the fact that news sites regularly edit published articles already, from fixed typos to corrections to large re-editings.
This isn't about a small percentage of stories, it's not about tone, it's the fact that if the NYT ever did this even once with the intention to truly "alter the past" it would be a major scandal.
And obviously things like corrections or taking down libelous content aren't included.
So no, I'm not constructing any kind of straw man here. I'm saying that the threat of subtly nefariously "altering the past" isn't realistic because it would be caught and exposed and there's no financial motivation to do it in the first place.
This is already happening without generative AI, and this new stuff is only going to speed things up exponentially.
The difference is that the floodgates are being opened.
We have lots of tools to fight spam, and there's no reason to believe they won't continue to evolve and work well.
Might not be possible on platforms - only if it's posted on a trusted domain.
That's the entire point of having trusted sources. Regular people can post whatever fake things they want on their own accounts; they can't post to the BBC's YouTube channel or to the NYT's website.
We haven't been able to generate 1,000 different forged variants of the same speech in a day before.
> We have certain publications and sources we trust, whether they're the NYT or a respected industry blog.
We can't even be sure that most of these aren't changing old stories, unless we notice and check archive.org, and they haven't had them deleted from the archive. The NYT has blockchain verification, but the reason nobody else does is because no one else wants to. They want to be free to change old stories.
You're wildly assuming a motive with zero evidence.
No, the reason companies aren't building blockchain verification of their stories is simply because it's expensive and complicated to do, for literally zero commercial benefit.
Archive.org already will prove any difference to you, and it's much easier to use/verify than any blockchain technology.
I'm saying this is going to become increasingly important fast, and we may miss the window where now almost everything not properly indexed by a large media organization is invalidated as there is no way to verify it.
I have a picture of Frank Sinatra at Disney World riding the tea cups. Who is the Frank Sinatra media authority that can tell me if this ever happened or not? A very small example to extrapolate from. It's going to get worse when everyone can create audio/video/pictures/text of anything they can dream.
The past may very well become a fictional dream, mythology, most of it impossible to verify.
Yeah this is not true. Sota Text, Image generation is well above average baselines. You can certainly generate professional level art on Midjourney
*my favorite is always the nightclub scene that goes real quiet when the actors act using their voices (which are real, but may be dubbed in afterwards).
We'll likely be able to verify whether an entity is a real human, using some kind of "proof of humanity" system.
We will have cameras/mics with private keys built-in. The content can be signed as it's produced. But in this case, what's stopping me from recording a fake recording?
Maybe it's a non-issue. We used text to record history and we've been able to manipulate that since, well, forever.
That'll make musicians happy with big tech as well, just like artists are. *sigh*
Then use the user’s action to iteratively refine your classification, until you end up with something tailor-made.
I suppose they might try, anyway.
This is precisely the opposite of the context I was remarking on.
> Mitigations: Vocals have been removed from the data source using corresponding tags, and then using a state-of-the-art music source separation method, namely using the open source Hybrid Transformer for Music Source Separation (HT-Demucs).
> Limitations: The model is not able to generate realistic vocals.
(https://github.com/facebookresearch/audiocraft/blob/main/mod...)
I suspect this was a combination of playing it safe and that the model isn't well architected to reproduce meaningful vocals.
Seriously, you couldn't sell this output for a free mobile clicker game.
Bruh, music is subjective as hell, and I can already tell I hate this song.
And if we now find ourselves inside this kind of world of illusions created by an alien intelligence that we don’t understand, but it understands us, this is a kind of spiritual enslavement that we won’t be able to break out of because it understands us. It understands how to manipulate us, but we don’t understand what is behind this screen of stories and images and songs."
-Yuval Noah Harari
If instead you consider that this new form of 'alien' intelligence is actually a descendant of human intelligence, that we are raising a new species which will inherit what humans have built (ideally only the good parts) and then improve upon it further..
It may sound grandiose, but that perspective changes everything.
If they just keep their models, people won't be interested and will build over ChatGPT or Bard.
That said, there are a ton of “look at this cool thing out research team did” and then you never hear about them again things from Google. They even built a music generator that was closed to the public until recently.
https://blog.google/technology/ai/musiclm-google-ai-test-kit...
They are trying to kill the market before they get left out.
If you're a company and wanted to integrate an LLM into your product if the choice is between several equally good models, but one is free and open-source which would you pick?
Aside from keeping competition at bay, this move also gives Meta leverage because ecosystems are now being built around their projects. If these models see wide-scale adoption they could later launch AudioCraft+ as a licensed version with some extra features for example.
Alternatively, they might offer support or hosting for their open source projects.
Right now though I think the primary benefit of these open sourced models is to attract talent. If Meta is seen as one of the leaders in AI then researchers will want to work for them simply for the prestige.
Arguably one of the reasons Meta has been behind so many awesome projects like PyTorch and React over the last decade was because they were seen as the cool place for recently graduated, but talented software engineers to work in ~2010.
I wondered though, generative AI is hurling us into a world where we'll need more mechanisms to sort real from fake, provenance will play a large part, and meta's platforms could be part of the answer. i.e. content linked to actual verifiable people.
He touches on it briefly in this podcast episode: https://www.therobotbrains.ai/who-is-yann-lecun
Curious if I’m alone in that.
(At the bottom https://audiocraft.metademolab.com/musicgen.html)
For what it’s worth though, the voice based examples sound dramatically better with MBD
Saying that engineers don't understand the arts is a bit of a trite generalization, but reading the way Meta markets these "music making" contraptions is really cringe inducing. Have you ever, at least, listened to some music?
/edit On a more serious node. I already see the 24/7 lofi girl streaming generated music. The sample[1] on lofi sounds pretty good.
[1]https://dl.fbaipublicfiles.com/audiocraft/webpage/public/ass... "Lofi slow bpm electro chill with organic samples"
Samplers have been around since the 70s.
Can I haz full version of Bach + `An energetic hip-hop music piece, with synth sounds and strong bass. There is a rhythmic hi-hat patten in the drums.` please?
(https://dl.fbaipublicfiles.com/audiocraft/webpage/public/ass...) ?
It's an interesting thought experiment, though. I can imagine that "environmental audio" companies like Muzak have about 5 years left before they either adapt or die. What other kinds of companies are in trouble?
Or alternatively, if the labels are not stupid, they'll negotiate for a higher price per listen (or similar), as they are still as essential to the service as before.
There doesn't appear to be any new CLI executables installed, and the documentation links to an API but there's no clues on how to actually process a prompt.
What am I missing? Alternatively, I wouldn't mind using it in a Notebook but so far this thread doesn't link to anything so ambitious (yet?)
python demos/musicgen_app.py
Otherwise you can check the jupyter notebooks in the same folder.I carefully went through the output generated by the "pip install -U audiocraft" command, and there were no clues provided.
Disclosure: I am not a Python developer, so I apologize if this is a master-of-the-obvious question for Python folks. However, if there was ever a scenario where a line or two of post-install notes would be useful, it's stuff like this.
I feel like Python folks are on average terrible at distributing software. So many projects have some python script to install the dependencies, still assume you use conda, or don't bother to specify the dependencies versions. Thankfully it's often the same patterns and after some time you understand what to do based on the error messages. But I wish they could use something like NPM or Cargo. Even something like Maven would be an improvement.
Disregarding the tip above, determining where the library was installed requires a bit of context, for example your platform (Windows versus UNIX) and the fact that newer pip releases default to "pip install --user" when not running with super-user privileges, whereas older pip releases did not default to "pip install --user".
Assuming you are using Linux and using an up-to-date pip release and you ran the "pip install -U audiocraft" command without super-user privileges, then the library was most likely installed in ~/.local/lib/pythonX.Y/site-packages (where X.Y is the version of Python that was used by the pip command you ran).
This setup takes 5 minutes:
- Mac Studio M1 Max 64GB memory
- running musicgen_app.py
- model: facebook/musicgen-medium
- duration: 10s >.if torch.cuda.device_count():
>. device = 'cuda'
>. else:
>. device = 'cpu'
So pytorch will fall back to CPU on a Apple Silicon. Ideally it would use Metal for acceleration (MPS) instead of just plain 'CPU', but if you replace CPU with MPS you'll probably run into a few bugs due to various Autocast errors and I think some other incompatibility with Pytorch 2.0.At least that is what I ran into last time I tried to speed this up on an M1. It's possible there are fixes.
I’ll have to check again, but I remember AFAICT my hardware wasn’t getting saturated, so maybe there’s headroom for mac cpu performance. And of course in the meantime I’ll be refreshing the ggml github every day
I know training data would be much more harder to get, (notwithstanding legal ramifications), but I think that creating structured, procedural data will be much more interesting than just the final, "raw" output!
While the datasets used for training AudioGen aren't available, is there any kind of list where one can review the tags or descriptions of the sounds on which the model was trained? Otherwise how do you know what kinds of sounds you can reasonably expect AudioGen to be capable of generating? And what happens if you request a sound which is too obscure or something not found in the dataset?
What are AudioGen's capabilities regarding spatial positioning? First example: can it generate a siren that starts in front and moves left to right and complete a full circle around the listener? Second example: can it do the same siren but on the Y axis, so it start at the front, it goes over the listener and then it goes under them to complete the circle?
How many jobs would this thing take away? One of the biggest time consuming in any video production is post production audio including background music, audio, Foley etc. This will automate almost of it!
I can't really understand this. I'm a DJ and a huge music nerd, and I spend a lot of time every week discovering new music from the past 100 years and all over the world, and I'm constantly struck by _how much of it there is_. I've spent weeks just digging through psych-funk records from West Africa from the 1970s.
How can you have the impression we're so desperate for more music that we need computer programs to generate it for us?
I don’t think this replaces “100% organic, human-made” music, though. I think there’ll always be a reason to listen to music made by other people. But I think this changes the landscape of how and why people create music to begin with. It certainly will devalue existing music, since everyone has something they may prefer that they can generate instantly.
I think generative AI is a terrible technology for artists who want to make money from their art, but in my personal opinion, I strive for a world where art isn’t a transaction, but a gift of human expression and connection. A world where art is appreciated for the emotion, stories, and ideas it conveys rather than the monetary value it holds. Generative AI might disrupt the traditional economic models in the art world, but it also opens up new opportunities for creative exploration and personal expression. It’s a challenging evolution, but one that could potentially democratize art, making it more accessible and personal than ever before! Bring on the Renaissance: Part 2!
In a world where nobody is compensated for their art, the only people making art will be the ones privileged enough to have the means to do so for free. I don't see how this leads to "Renaissance: Part 2."
this isn't democratizing art, and i would argue it has nothing to do with art. it is giving us an endless faucet of content, but not art.
Recorded music is the worst thing that happened to music.
I came to understand a very large portion of the population just wants content, any type, any quality, to fill the void. They'll consume anything as long as it's new. Content to fill the empty vessels we became. Just look around, mainstream music, movies, podcasts, news, it's mostly mediocre, but it goes real fast, you get new mediocrity delivered every day
As an aside, a lot of musicians seem to dislike this kind of technology, but I never saw music as a competition. I don't care if some inexperienced kid is generating bangers from his bedroom even though he can't play a single instrument. It's just something else to listen to. I write music for me.
Could you generate a rhythm track? Ideally you could make songs one track at a time, by giving it a mix of the previous tracks and asking it to make another track for an instrument. Or, give it a track and ask it to do some kind of effect on it.
Another interesting use might be generating sound samples for a sampled instrument.
They both work in varying degrees of success. Audiocraft on github, issues or discussion sections have a lot of questions answered.
I don't know if audiocraft_plus incorporates all three modalities of the release, MusicGen, AudioGen, and EnCodec. It uses MusicGen for sure, all four models, small, medium, large and melody.
I haven't looked closely to this release, is the audiocraft page on github different than facebookresearch/audiocraft? The other two modalities may be new, AudioGen, and EnCodec, but i was under the impression that they changed the license to full open source, and that was that.
So yes not a literal repeat, but no not enough to rate as a "full song", IMO.
Plus it is not actually song at all, note. Singing would particularly expose its repeat-and-vary trick.
What about the ring tone, busy tone, disconnected tone for any country over time. 2600 vibes (pun)
The samples included in the press release are quite impressive to my ears, but the other samples (especially from AudioGen) have a hint of artificially.
As usual the music is quite repetitive, but I'm looking forward to tools that simplify changing the prompt whilst it generates over a window. I can only imagine the consequences for royalty free music.
Edit: the "Text-to-music generation with diffusion-based EnCodec" samples are quite impressive.
"From text to audio with ease"
I hoped for a second we would get a good quality model to do text to speech - damn, I guess it's back to bruteforcing bark.ai or waiting for tortoise (or more realistically, just paying elevenlabs)
The code is MIT, though.
If they can’t for that reason alone, then the model is a mechanical copy of the training set, which may be subject to a (compilation) copyright, and a mechanical copy of a copyright-protected work is still subject to the copyright of the thing of which it is a copy.
OTOH, the choices made beyond the training set and algorithm in any particular training may be sufficient creative input to make it a distinct work with its own copyright, or there may be some other basis for them not being copyright protected. But the mechanical process one alone just moves the point of copyright on the outcome, it doesn’t eliminate it.
In this instance, it's very likely that the process is sufficiently transformative. A set of model weights look nothing like the Mona Lisa, nor can they be directly transformed back into it. What it is NOT is the product of a creative process on the part of a human, and is thus ineligible for copyright.
It is as though we are able to distill meaning using an automatic process. Copyright doesn't protect meaning, only expression, and it only protects expressions that were generated by humans.
The network itself might be patentable if it isn't obvious.
If a network was copyrighted and it was found that the function of the network was inseparable from its expression, it would no longer be eligible for copyright on that basis.
I think people often misunderstand the application of copyright to GPL code in this way.
Now, will any of this stop people from CLAIMING copyright? No. It'll have to be fought out in courts.
Like for real.
The last bastion of human creativity is about to be defeated.
But clever devs can study it to make better software and pressure them for a better license on the next release. Worked for LLaMA 2
Audio is more similar to language than images because of more stronger time dependency.
The paper says the critical step they took for making diffusion model work for audio was splitting the frequency bands and applying diffusion separately to the bands (because full band model had limitations due to poor modeling of correlations between low frequency and high frequency features).
I think something could be done on text side as well.
So, for text: a) what is the equivalent of a small, noisy step? and b) what is the equivalent of a uniform gaussian in language space?
If you can solve a and b, you can make diffusion work for text, but there hasn't been any significant progress there afaik.
It will be great when there's eventually something open that competes with the closed models out there.
We might know more had it generated a song as you said, but in fact it generated only an instrumental.
It is not to "own more rights" to shit. Sorry, but no one actually owns what they create, that's the point of creation. And if you take issue with the term create, that only reinforces my point. We're all influence machines, input and output, the future should not be about preserving some peoples rights to limit our collective advancements over their personal wants. Tough shit.
Our overlords didn't get the memo
If we can off load the mundane (survival) aspect of our pattern recognition engine, then maybe we can use those cycles in lofty pursuits - this is the victorian fallacy.
-
Break everything down - its all patterns all the way, and how we process them... we are letting AI take on an aspect of ourself (pattern recog ;; tokens;; and prediction)
That is, if applied to self-preservation, is the essence of sentience.
(I think! therefore, I am, and I will prevent you from making me NOT)
I mean everyone who writes right now, into this HN thread, owns copyright to someone for the letters, amirite? That someone is me. Anyway, long story short, copyright was always a pretty ridiculous idea, alongside with patents of course, but it is only right now, with programs that can mimic writing style, painting style, speech style etc, that is obvious to everyone.
As a side note, there was a Greek private torrent tracker, blue-whitegt, which today would be a serious competitor to American companies like netflix or youtube, but it was shut down because, surprise surprise, there were some copyright issues, despite the site being a really quality service, a paid service of course. Blue-whitegt today would be a 10 billion to 100 billion company, instead all the profits aggregated to American companies.
When it comes to copyright, very soon everyone on the planet, will have monetizable torrent seeding. All of these copyright chickens, are coming home to roost!
https://en.wikipedia.org/wiki/Archaic_Greek_alphabets#Euboea...
Agreed!
> the future should not be about preserving some peoples rights to limit our collective advancements over their personal wants.
Systems that don't allow people to extract value from the hard work they put into collective advancement do not seem to lead to collective advancement over time and at scale. Incentives matter. No one's going to spend all day making candy and put it in the "free candy" bowl when that one asshole kid down the street just takes all of the candy out of the bowl every single day.
At small scales (i.e. relatively few participants with a relatively high number of interactions between them) then informal systems of reciprocity and reputation are sufficient to disincentivize bad actors.
At large scales where many interactions are one-off or anonymous, you need other incentives for good-faith participation. There's a reason you don't need a bouncer when you have a few friends over for drinks, but you do if you open a bar.
On the other hand this is what researchers do all day every day. PhDs and professors work for the common good and get barely any pay in return. Maybe the future model in art and music is more like the academic researcher.
Artists don't see a single red cent from their work being sucked up into some AI content blender. Their work is being taken and used-- often in service of others making a profit-- and they receive nothing. Not even credit.
Edit: Well, they don't receive nothing-- they get a bunch of people telling them they're selfish jerks for wanting to support themselves with their work.
This is how it has always been, and fundamental to the economics of art. Things people are willing to do regardless of financial compensation rarely pay well.
Deviantart in the 00s was a huge repository of art people were mostly making for free. Some people got lucky and turned that into a full time occupation, but the vast majority didn't.
None of the people whose art was sucked up by these machines had any idea that this would be integrated into for-profit tools and used against them in the market place, and almost certainly would not have consented if they had known that. The ones with copyrighted images didn't even give legal consent. The fact that Getty's logo showed up in red carpet images is a symptom of a problem that obviously went well beyond Getty, but there's no way an independent artist could prove it.
Furthermore, if you think artists + luck = commercial art you're completely 100% wrong. Most art school graduates don't go into fine art for self-expression with some lucky individuals matriculating into careers-- they go for job training. Go look at the degree programs for any art school-- almost all of them translate into a full-time commercial career immediately out of school the same way a STEM degree does. Concept artists, illustrators, character artists, environmental artists, graphic desginers, VFX artists, animators, cinematographers, choreographers, commercial musicians and composers, photographers... these people don't just spring up out of the fine art world. That is their career. I know because I am one.
Your ethics-dodging free market tech libertarian garbage holds no sway with me, so you might as well just save yourself the keystrokes.
AI allows more people to be artists without the skill barrier, and that is a social good.
I'm sorry that the mechanistic portions of your art career are rapidly losing economic value, but I think free tools like StableDiffusion (and the better tools that are coming in the future) should be available to every child (and adult) in the world. And the world will be a better place as a result.
The utilitarian argument is only philosophically defensible in the street car scenario; when people are deliberately pulling the levers unprompted and it could have been done ethically but it wasn't because it was just to darned inconvenient and/or expensive, the greater good argument doesn't work. It's the same argument people used to defend the Tuskegee experiment and its ilk... And roman public slaves. If we're willingly throwing people into the spinning wheels of progress because it will be wonderfully convenient and neat for non-artists to use other people's skills instead, there is just no ethical defense. Knocking the stool of from under someone using a tool they built and you didn't even have permission to use is just not ethical. You and many others end up falling back on right-wing platitudes about the free market.
I'm actually a technical artist— my skills are now dramatically more valuable. That doesn't make it any more ethical even if it works on my favor.
I was a developer for 10 years, and have seen this unadulterated hubris in nearly every group of developers I've ever encountered— when you find yourself explaining someone's job and industry to them, you might want to stop and ask yourself... "Am I pulling a Dunning-Krueger right now?"
Here is where I think you're wrong: Everyone is a potential artist.
Many never have the opportunity because of economic reasons: some people don't have the time to cultivate the mechanical skills necessary to render their vision.
All artists pull from the work of others constantly. That is 'using other people's skills'. It is absolutely privileged to have the opportunity to devote time cultivating art production skills (which is usually done by studying and replicating the process of the artists that came before them).
If someone is born to a single mother, grows up having to raise their siblings, then has to immediately start working to help support their family, they don't have the privilege or time to explore art production skills. Technology has been making art production more accessible for hundreds of years, and that's a clear net good. I want everyone to be able to make art if they want to.
It up to the artist if they want to manifest their vision without creating the pigments and brushes themselves, or using a drawing tablet, or using generative AI tools - it all can be art that is meaningful to the person who creates it and more art in the world is a good thing.
Commercial art isn't a very good representation of the artists soul: they are constrained by market forces or their patron rather than producing what their own heart desires. We should all be happy to sacrifice commercial art jobs if it means that every person in the world gains the ability to render the art that comes from their soul.
Being a professional artist isn't fucking magic. It's not winning the lottery or even getting drafted for a pro sports team. It's a large group of regular fucking careers just like any other but creativity makes up a larger percentage of their professional cognitive toolkit. I know a workaday oil painter who's neither well-known nor rich but went to college to learn how to be an oil painter and that's how he comfortably pays his bills. He paints seaside landscapes, wealthy people buy them to put in their vacation homes, and that's how he makes his money. He was in college right along side with graphic designers, animators, architects, product designers, user interface designers, and all manor of other professional artists. Just like any other non-licensed white collar profession, people who didn't go to school are in the business, but it's harder. I also know a professional comic book artist, many animators, sculptors that work in product development, and plenty of other professional artists that built normal careers like any other white collar professional. The fact that not everybody can go to art school to build an art career doesn't change the disposition of professional artists any more than not everybody being able to afford to go to medical school affects the disposition of doctors, and having "the soul of a healer" doesn't really enter the goddamned equation, does it?
Whether you're being deliberately obtuse or are painfully ignorant about something you've got a lot of strong, baseless opinions about, you're obviously not going to cut the bullshit and be honest with yourself. The reason you have to delve into all of these pseudo-philosophical mental gymnastics is because you're wrong, but you really don't want to be, so you're trying to construct a reality in which you're right. That's not how reality works. Using people's work without their permission to make a for-profit system to compete with them in their professional marketplace is not moral no matter how big of a castle of bullshit logic you build around it.
Bye bye. I'm going to let you hang out in your little land of make-believe by yourself.
That's a false dichotomy. Just because the economic issue and cultural impact are separate doesn't mean they're mutually exclusive. If a mugger stopped taking people's money and instead just walked around making people afraid for their lives by intimidating them or beating them up, that would still be immoral.
Whether people consider it immoral has no bearing on whether or not the cultural and economic issues are mutually exclusive-- they're not. You can't say that the argument is only about the economic issue simply because there is an important economic argument. It just doesn't make sense.
Academia is a carefully constructed system whose incentive structure is based on highly visible explicitly measured citations and reputation.
People aren't generally just trying to maximize wealth. They're trying to maximize their sense of personal value, which tends to be a combination of wealth, autonomy, and social prestige. Academics (and some creative fields) tend to be biased towards those who prioritize prestige over wealth.
Software is an infinite candy bowl. Taking candy out of the bowl does not take any candy from the person who made it.
Imagine if this was the physical world, and you had a machine that could end world hunger. You could copy food like you could copy and paste information on a computer. Imagine someone who would keep that machine to themself, out of a sense of entitlement to make a few bucks. Any person with any sense of morality can see the obvious problem with that.
More in theory than in practice. Ask any open-source maintainer how much running a popular project is unlike putting out an infinite candy bowl and then going on with your life.
Software is not a finite candy bowl. It is also not an infinite candy bowl. It's not like physical goods at all, not even like physical goods that can be magically cloned. It's just different, entirely.
The incentive and value structures around data creation and use just can't be directly mapped to physical goods. You have to look at them as they actually are and understand them directly, not by way of analogies.
Why do people make software and give it out for free? Is it purely from the joy of creation? Sure, that's part of it. The desire to make the world better? Probably some of that too. Are those forces enough to explain all open source contribution?
Definitely not. Here's one quick way to tell: Ask how many open source maintainers would be happy if someone else were to clone their open source project, rename it, claim that they had invented it, and have that clone completely overshadow and eradicate their original creation?
If the goal was purely altruistic, the original creator wouldn't mind. More candy in the infinite candy bowl, right?
But, in practice, many open source maintainers strongly oppose that. There is a strong culture of attribution in open source, largely because there is a compensation scheme built into creating free software: prestige. One of the main incentives that encourages maintainers to slave away day after day is the social cachet of being known as the cool person who made this popular thing.
> Imagine if this was the physical world, and you had a machine that could end world hunger. You could copy food like you could copy and paste information on a computer.
Analogies are generally bad tools for real understanding, but let's go with this. Let's say this machine took fifty years of someone's life to invent, toiling away in obscurity. Basically, an entire working career spent only on this invention with nothing else to show for their adult life.
If, at the end, no one would ever know it was you who invented it, how many people would be willing to sequester themselves in that dark laboratory and make that sacrifice?
Candy can be consumed to depletion. Art gets richer the more it is consumed.
A much better analogy would be having a sculpture in your front yard. The idea that a kid would be an asshole for appreciating the sculpture too much is obviously laughable. People choose decorate their yards for the status having an attractive yard brings, without the expectation of profit from it.
So is the idea of real property being something that someone owns.
If nobody is beholden to any job or duty, and the machines do everything, who is to say I don't want to make every machine on earth dance in a flash mob? I cannot do that, because it would require other people to halt their use of the machines. Abundance is a false promise and one we should be quick to shoot down lest we surrender our future rights to the ones advertising it.
Removing the worth of people in their jobs removes their leverage in the constant resource allocation negotiation in the economy. Given that we just witnessed Elon Musk spend 20,000 average American lifetime earnings worth of wages just to be the new dictator of a social media company, I'm not sure that I want those negotiations to take place only among the giga-rich.
Thankfully your opinion is an extreme opinion and will never come to pass. Tough shit indeed. :)
----
I really like hackernews but recently I've been seeing a plethora of "your rights don't matter, you own nothing" bs spreading around.
> I've been seeing a plethora of "your rights don't matter, you own nothing" bs spreading around.
It's not a new sentiment. Plenty of countries worldwide ignore US copyright, if they all got sanctioned then America wouldn't have have electricity.
If you author a copy of your content in a digital medium, you should be prepared for that content to be redistributed against your will, infinitely. It's not the nature of humanity, it's the nature of the digital format.
That is fine.
>If it makes me happy to write my name on the cover and pretend I wrote "Robert Frost's Poetry Collection" then so be it.
You can do whatever you want in your own house. No one is trying to dictate terms of what you do in your own private time in your own domicile. The problem occurs when you try to profit from my work, by pretending you did the work, and selling it to the public pretending to own and create the work which you have not.
>I can even sell that adulterated copy under the First Sale doctrine.
However, the idea of "You can add a chapter to this, call it your own, and sell it" and I won't legally come after you is absurd. and if the results of said legal pursuit result in you being bankrupt...what was the phrase you used, "then so be it."
---
Alternatively, you could write your own fictionalized work and sell it. But that requires more work than copying what I wrote, calling it yours, and selling it, doesn't it?
The only part of this that seems legally problematic is copying and redistributing the part that you wrote. If I wanted to buy 10,000 copies of A Game of Thrones so I could staple my fanfiction to the back and sell it on Etsy, there is no legal precedent suggesting I could be stopped. It is a lawful transformation of a legally licensed product, payed for in-full and redistributed in accordance with it's individual license. Absent any extenuating contracts between the owner and seller, I don't see what legal ground you would have to stand on.
> and if the results of said legal pursuit result in you being bankrupt...what was the phrase you used, "then so be it."
"if" indeed. Let's check in with the Author's Guild and see how this fight is going: https://www.copyright.gov/fair-use/summaries/authorsguild-go...
They are with you on the advance, and that in the long term science and the useful arts can't be owned. But to achieve that long-term goal, they saw it as valuable to give people temporary rights to align those "personal wants" with "our collective advancements".
Don’t want to pay for content ? Well we have “solved that”…
How is pressing a key on a piano different from pressing a key on an electronic piano?