MusicGen: Simple and controllable music generation
ai.honu.io
ai.honu.io
Music producer here with an honest question to those saying "this will provide me with a simple soundtrack/background music for $PROJECT"
Have any of you checked out / made offers on music production subreddits? Or other music subreddits? various music production discords? Elsewhere on the internet?
If so, could you say what your experience has been?
I ask because the music production scene is like...ridiculously saturated, and it's almost a meme in the producer community how hard it is to make even a buck producing. I suspect that there are a significant number of producers who would be happy to take your "prompt" for a small fee. Yes, I understand 1) free and 2) immediate is convenient, but isn't 1) relatively inexpensive and 2) whatever advantage intent in construction gives good too?
I'm willing to admit that I'm missing something here, but I'd love it if someone could enlighten me.
While I'm asking follow ups, to all the folks who love digging for new music so much that they're considering turning to prompting AIs, I'd be seriously surprised if you've really checked out all the stuff that is coming out from new producers (again, reddit, soundcloud, etc). Another meme in the producer community is how one spends hundreds/thousands of hours perfecting ones craft, and dozens of hours working on a track, only for that track to get like 5 plays on soundcloud and negligible engagement elsewhere. Are music consumers really that desperate for new tunes? Frankly a lot of us just aren't seeing it....
I doubt I would turn to AI much for anything other than background noise while focusing on work. In fact, that sounds like a perfect use case for me. "Dear GPT, please compose a four-on-the-floor downtempo progressive track with soft pads, no vocals, and zero goddamned fake vinyl noise that runs for two hours straight..."
Yep. this is why I don't feel like AI used in this manner moves the needle for music: people only actively listen to the best 0.1% of music anyway. The ability to create music that is firmly in the other 99.9%, as this stuff very clearly is, just means that the ocean of mediocrity has more water dumped into it.
To the extent that sounding like everything else is a problem, how is ML generated music not going to have it?
And in general this isn't going to be a qualitative improvement in experience. ML algorithms for recommendation are searching the preference space in much the same way ML generation would, they're just doing it over existing stuff. If you really find 99.9% of existing material boring you're probably going to find a similar order of generated material boring.
Though I suspect 99.9% is hyperbole. My rate of "this is listenable and interesting and I'd like to come back " on Soundcloud is better than 1 in 25 on the worst day and better than 1 in a dozen on most, and the rate is often north of 1 in 6 for curated platforms like Pandora. It's never been easier to discover good new music to listen to with not much in the way of effort.
AI generated art has explored all sorts of weird spaces that few humans have touched.
It's not difficult to make computers create unusual, original, bizarre work. The difficulty comes in making it both original and enjoyable/interesting.
Also consider that AI-generated music is often going to actually be a collaboration between a human and an AI. The human will be acting at least as a curator, because not everything created by AI is going to be pleasing, so some selection and catering to human taste will be required.
We may also disagree on how much good stuff there is coming out, but I agree there is a lot of noise.
Tackling the latter would likely exceed the entire effort I spent writing my little hobby game in the first place. I don't think it's even close; it was never a serious consideration. Some of these games I write in a single sitting. I do my best to piece background music together using a chord progression app, descriptions of keys and the notes they contain from Google, and premade drum loops and instrument samples. It comes out worse than if a real musician had made it, but getting a real musician was never really an option.
It's the same for the art. I don't have the time or money to pay an artist. They deserve to be paid fairly for their work just like musicians, but I don't have it and it's just a stupid hobby game. But even stupid games need art and music. So, homemade programmer art and music it is. The availability of better tools to help non-musicians hack something together is greatly appreciated. I haven't tried any AI stuff yet but I will next time.
If your project is small and free, you're not going to land The Eurythmics. But all those people posting their music online hoping to get noticed? Emailing them, even a cold call, immediately tells them you've listened to their stuff and you like it. Honesty is the best approach.
I think OP is onto something.
Edit-
> but not be 100% sure if I own the rights or not).
That's also really easy: stipulate it in writing. Preferably a proper contract but an email agreement is defensible too (IANAL).
When it's just me and some apps, I'm writing the background track in an evening after coding the game earlier in the day. If I bring someone else in, now I'm writing contracts (something I'm completely unprepared to do correctly myself, as a non-lawyer). It's too big of a jump for a one-day zero-budget hobby game that isn't very good and only my friends will play. Not when I can quickly cook up something myself using readily available tools.
For a more serious project with a budget, absolutely you find a professional producer, just like you get professional coders and artists. But this isn't that.
I've offered this to like 15 people. Some respond in utter confusion and blow me off. Most don't even respond.
This isn't 2005. The vast majority of people, especially "music artists" are not corresponding over email and are not exactly professionals either.
You are missing the point, most programmers are introverts and absolutely deplore doing cold calls and cold emails.
... this is almost TOO convenient.
- Find me, which is easy, just become an amateur developer yourself and scour the various places I frequent
- Think of and propose in great detail what you want. We will go back and forth over this, over the course of several days/weeks. Bonus points if we don’t speak the same language.
- Sign some form of agreement, really easy, just read this 5 page document and maybe hire a lawyer if you are unsure. All very easy.
- Fork over the cash
- Get deliverables in a few weeks, hopefully.
Now when you compare that with just firing up some website/app and getting on with your work, is that really better? I’m not seeing why you would just go to gmail.com if I could have just made you a very nice, very special email reader.
I have perfected my craft over thousands of hours you know. You should pay us the respect we deserve.
—
Seriously: cranking out tunes through some prompting vs hiring people through shady channels like reddit? Are you serious?
You give an analogy of software. I suppose I feel that art differs from this in some ways. E.g. originality, creativeness.
Reading many of these responses I'm gathering that perhaps I have too strong a notion of what kind of quality of music might be in demand. While the loops on the original article are impressive given their generative nature, I suppose I felt that there may be a demand for something more (Better sound design, more long term structure), but maybe I'm naive.
But thanks for giving me your perspective.
You hit a nerve because software development is also art and highly creative. I have given a significant part of my life - basically my youth - to it and I feel “creatives” think they are somehow special and that their work is fundamentally different and I don’t think it is.
It’s just that we have been cornered earlier than you guys. My skills are now only profitable as boring building blocks in corporate settings because nobody else will pay for proper work. Everybody expects easy access for free instantly to whatever digital service they can get their grubby hands on. If I talk about “craftsmanship” I get laughed out the room. Nobody gives a shit.
Now I’m like, yeah guys, that’s how it feels to have your skills commoditized. Deal with it. That’s kind of childish though.
The problem I feel is that I have an expectation of being able to front the cost of engaging someone to work on a project with me.
Working out navigating a working relationship on a smaller project seems fraught with issues.
I'm rarely inclined to spend dozens of hours listening to soundcloud when I have other things to work on.
I mean yes people create interesting music, perhaps it's a search problem? Knowing someone creates the kinds of music I'm interested in would help. But as someone making things, I'm trying to find someone who I can collaborate with who has an overlapping interest in what I make. Solving for that is not straightforward.
I've had much more luck with graphical art than music.
So yes, even though these systems are fundamentally worse, I can at least "collaborate" with them on producing something. Going from zero to one can be enough.
For music we could present such an image but it would then suggest I'd argue much more possibilities. We could narrow down by genre you would suppose but even then there are too many possibilities: genre's are not as strong categories as are the stylized "era's" of visual art, I would also claim. Moreover, we can "port" a fundamental structure like a melody over all sorts of strains of music. In visual art, any motif is bound to be changed depending on the era and the style we'd put it in, that is, I think that in music, there are elements that are stronger in visual arts and elements that are weaker in music, and vice-versa, with regard to a description we could give in English. It's probably more natural and more possible to ask about what a sort visual representation should be than what a piece of sound should be.
It's interesting how we can generate images I'd argue in stunning faithfulness to some prompts but we don't seem to be very close to the same standard, for some prompts, at generating music.
The guys in Infected Mushroom will have a field day with this stuff. Their whole thing is finding weird ways to create new sounds you never heard before.
Just another instrument, really.
If anyone knows anyone working on that, ping me. :)
When it's 3 am on a Saturday and I'm in the zone on a passion project, I'm not about to spend the rest of the weekend going back and forth with a music guy on Reddit.
I want a music robot that cranks out music on demand and responds to my every whim, and a real human being isn't going to want to fill that role no matter how impoverished they are.
Add to that, if I don't like something, want it tweaked, want something completely redone, or just flat out change my mind about some direction I provided later, I have to go back to them and negotiate a new contract, or find someone else to do the work. The costs add up over time, and there's an additional benefit to immediate feedback (or cost of delayed feedback, as anyone who has worked on a software project that takes forever to compile/check can attest)
I haven't used the music AI tools yet, but having played around with Dall-E a bit I can say that it's pretty enjoyable to be able to give direction, and bound it, then roll the dice and see how things turn out. I definitely feel some ownership of, and pride in, the resulting creation
All things equal, people are happy to support local businesses. The value prop here is far from equal.
the real meme is about how artists have always been grasping for financial respect in every market condition ever, and yet nothing has changed. people were never going to commission you, they were never going to book you. While they do appreciate the content. But for the few that would ever actually try to commission something, they encountered friction after friction after friction and collectively artists have been disinterested in solving. Because they're starving and preoccupied with fighting for scraps and modicums of respect at all.
The world’s has now solved many of these frictions.
The frictions were:
1 hoping they found the right artist to begin with
2 hoping that artist is reliable and has any work ethic or structure in their life
3 not bruising that artists ego in however communication style is preferred
4 dealing with how completely segregated many artists are from contract negotiations and any aspect of the business world, but needing to secure rights properly
5 ego in securing rights properly without the artist overplaying their hand
6 waiting for the commission
7 revisions
8 circle back to 1
9 if you ever get past part 8, you have the issue of whether your new license can be used in an unforeseen way and medium in the future
getting burned in altruistic commissions of living artists is simply over now. all these frictions are solved with the free and immediate way.
> the real meme is about how artists have always been grasping for financial respect in every market condition ever, and yet nothing has changed. people were never going to commission you, they were never going to book you. While they do appreciate the content. But for the few that would ever actually try to commission something, they encountered friction after friction after friction and collectively artists have been disinterested in solving. Because they're starving and preoccupied with fighting for scraps and modicums of respect at all.
For anyone trying to make money off of music, they should have already been aware that most of the effort in making a living is the non-music work. Once your music reaches an acceptable level of quality it's more about finding and managing your fanbase, industry connections, getting booked at the right shows, promotion and marketing, maintaining professionalism, etc. than anything else. Which this particular AI doesn't help with.
An extreme example is Fred Again, who came out of nowhere and is now one of the biggest names in electronic music. His music isn't bad, but it's nothing revolutionary. As it turns out, though, he grew up in one of the richest neighborhoods in England, with Brian Eno as a neighbor, and went to the most expensive private school in London.
So no, AI music generation doesn't change anything here. It's similar to the startup mistake technical people make of focusing on picking the right tech stack instead of focusing on sales and finding product-market fit. The software/music is only about 10% of the challenge of making a successful business/career.
I did want to clarify that I was posting from an angle about those us who need music produced for our products, but were never going to commission it.
I think its important to understand that user story because a lot of artists don’t seem able to empathize with it. People are excited because they were never going to commission artists, and were also turned off from stock music licensing websites too.
Mine are and the same goes for most of the artist I indicate. The point wasn't that they were in it for the money, although many dream of being able to at least one day pay the rent with it (or maybe just groceries).
The rest of your response makes sense (although I think much of it could be said for all of hiring someone to do work). Anyway, thank you for providing your perspective.
As anti-social as it sounds, it's the conclusion that I've reached to after years of working with freelancers/contractors. I've contacted with >50 artists (>300 if "they send me a propose on Upworks" count) and worked with ~10 of them.
Don't get me wrong, I still choose human artists over Stable Diffusion. For now...
It's basically the same as with Midjourney. Before Midjourney I'd have to spend quite some time organizing with some human, explaining what I want, licensing terms, etc only to have to wait a significant amount of time for an image that I may not like.
With MidJourney for just a very small amount of money I can instantly get images that are exactly what I want, iterating extremely quickly. Just the fact that I don't have to deal with another human saves a massive amount of time.
TL;DR
1) Faster
2) Cheaper
3) Often closer to what you want, because you quickly iterate and can get hundreds of variations
I know the people who say that mean well, but it totally overlooks both how much the culture does (as you say) want to provide their art for projects to use it and create value together, and... reality. Shouting at everyone isn't the way to get them onboard, but shout they do, and it's one country in particular that seems to scream the most.
I'm in the UK and I can't walk down the street without tripping over producers, so maybe the way around the angsty people is finding them in person? Or... we just use AI. The robots solving our social issues is probably a thing.
I’m not going the AI-route. It’s frankly easier, more fun and it sounds better to compose the music myself, and if the project takes off and makes a dime, hire a pro to improve things later.
This is an insurmountable benefit. Literally the two most important things when it comes to me buying music at scale.
I made a mobile game a while back and composing and licensing music cost me $30,000 for a free to play game. That was the same as 6-months dev salary (devs in Belarus).
If I can save $30,000 and have zero delay, I’m just going to do that 100% of the time.
The only factor that a real musician can beat on is quality. But let me tell you, with zero marginal cost of production, quality will inevitably be better with generative.
I don’t think it’s going to displace a dedicated composer that gets the medium they are scoring for any time soon. But then that’s not what your comp was initially.
TLDR there are cases where “good enough” is going to be provided by generative music in the medium term. Unlikely for this to be anywhere adjacent to music connoisseurs.
Wow this is more than good enough to use for background music in video games, stores, commercials, etc..
You really could have super dynamic music in a video game for instance that changes based on the time of day, environment, situation, mood, etc.. all combined.
Combine it with a LLM DJ and you could get some fun radio stations.
Games can and do already do this, dynamic sequencing of music from a pool of stems has been common practice for a while. Maybe this could let you do it cheaper, and AI could go more granular by creating new stems on the fly, but the onus is still on the AI developers to show something which hits as hard as someone like Mick Gordons dynamic compositions.
Infinite variety is of little value if the infinite space is full of infinitely boring, uninspired content.
And no you can't do this already:
const musicSpeed = inFightSequence ? 'intense' : 'chill';
const musicPrompt = `drum and bass beat with ${musicSpeed} percussions`;
playMusic(musicPrompt);I spent like 6 hours yesterday playing with this. It's really cool, but not that good yet.
I'd say it's like the original stable diffusion (without any of the finetunes and improvements). Very cool, but not 100% there yet.
Hard disagree, and lack of copyright due to not being produced by a human becomes an issue for many video games, commercials, etc.
>You really could have super dynamic music in a video game for instance that changes based on the time of day, environment, situation, mood, etc.. all combined.
You don't need this IA for that at all.
>Combine it with a LLM DJ and you could get some fun radio stations.
You could also not, and you wouldn't know until it failed to produce anything interesting. A whole radio station filled with grocery store background music? oh wow I can't wait for the fun.
and you're not likely to make any money doing it, so what's the point aside from showing the human portion of music is missing in everything you suggested.
Suppose there is way to measure cardio beats or electricity spikes on the brain, and we configure the machine to generate music to increase cardio beats, or decrease them, or similarly increase electrical activity of the brain or decrease it. Then psychology might be deprecated, mood will be reduced to just a music channel.
[1]https://soundcloud.com/kwstas-pramatias/lounge-owl
[2]https://soundcloud.com/kwstas-pramatias/rock-glass-shatterin...
You are probably getting downvoted for your second paragraph which is a bit out there.
Yes of course they are the starting point, a good musician may take some samples and transform a music generation to a better song for sure. Some artists state that a painting is never complete, or a song is never complete. There is always room for innovation.
The prompts i used, referenced real songwriters, and the model seems to know their songs. The article does not prompt it that way. So i guess there may be a little bit of IP infringement, but we need that, only for the first bunch. Next models will be trained on the best generations of previous models.
Kind of like how the software is never complete :P
That can be achieved by putting music on, which speeds up the heart pulse. Usually hard rock, metal, thrash metal etc. In that case, the body starts sweating a lot, not matter the temperature. I combine that, with 5 simple exercises i do all day long which are important as well.
My point is that using music, someone can be in charge of his heart pulse. But my biggest complaint always was that these metal guys, are masters of the guitar, but other kinds of music have better taste in rhythm, in melody etc. Using programs like that we can evolve it a little bit, to be more pleasurable to listen.
I know about about binaural beats, i have tried to listen to different hertz for hours on end, they don't work in my opinion. At least in my case.
I remember vividly that this was very hyped in some circles around 2005 or thereabout, with wild claims that listening to some strange white noise for twenty minutes could induce full-blown psychedelic trips even in people with no psychedelic experience. I even tried a bunch of em, and the only clear effect was a mild headache. And I was naïve enough to think it might work back then, and yet there wasn't even really a placebo effect.
Oh, to be a fly on the wall in RIAA corporate offices…
Sans schadenfreude, I think this (depending on inference speed) could be perfect for dynamic content in games (including IRL games: LARP, escape rooms, table top games, etc.)
In all likelihood, they're ok with events. Games were never anywhere near their main revenue stream. Now the labour costs on what they're actually selling are dropping to zero. RIAA's future:
1) Use AI to fake a band.
2) Use AI to write music (maybe even lyrics). Don't really care if the AI is any good.
3) Distribute output widely, note that copyright still applies to the output.
4) Use media to generate hype (the critical step). This depends only on platform control/relations, and they have that.
5) Yea, other people could technically generate same quality dreck with AI, but it won't be (and legally can't be) exactly like the hyped dreck. Others can replicate nearly everything except the hype.
6) Since the costs are near zero just about every sale is pure profit.
Basically, since Music can be replicated, they'll sell hype and belonging to a fan group instead.
Making and marketing an AI band isn't even interesting. Someone will be doing it on twitch and youtube an anime vtuber ensemble before the RIAA even figure out any portion of it. The media hype is because of celebrity, and AI generated stuff can't be celebrity.
With sufficient social network campaigns, media brib^W relations, and paid influencers we can get anything to be a celebrity, whether it's a paid actor or an AI avatar. That's the special step that not quite anyone can do. Plenty of K-Pop is already not that different...
Can you elaborate on this point? I don't think this has been established to be the case yet.
So even if the output is technically made entirely by LLM, they'd find a way to slightly tweak the process or even the law so it applies. Someone will do a trivial low-pass filter and then claim copyright. At worst they'd find a flunky to say they 'wrote' the music.
Those two can still rely on patents and trade secrets in a way that RIAA can't. (At least, the RIAA can't rely on those as far as I can see, but what do I know…)
https://blabbermouth.net/news/anvils-lips-on-bands-using-bac...
I’m very impressed.
By contrast OpenAI's JukeBox, which is maybe two years old now, really comes up with fascinating ideas.
Mixing genres do not really work and the model doesn’t seem to be trained on band names. However it does perform well to create music using existing styles.
I generated some Eurovision crap and minimalist techno that were very much believable. But mixing death metal with lofi ambient isn’t the best, nor the epic progressive rock guitar solo I asked.
I think the examples on the website are cherry picked but with some experience in prompt engineering and many tentatives, it should be possible to generate great samples.
It’s also excellent at generating boards of Canada like music. The audio artéfacts, the low fidelity, the weird sounds, the detuned synths, this model does that very well and it does sound great to me.
Thanks a lot to the authors.
Would you expect that to be good if a human did it?
The style transfer is the most interesting bit IMO, as you get a sense of how it hears the source examples.
For example, when transferring the opening to the Bach Toccata all the new samples miss out the same passing note (the fifth note in the sequence). To a human ear that note is important, and could easily have been incorporated into the new samples, but it seemingly doesn't activate enough neurons for MusicGen to care.
I mean like YESTERDAY I did not have this superpower to summon something as majestic as say https://fb.watch/l4ssOD40M4/ with a simple 'A quirky and skronky Aphex twin sample that just hits you'
Edits:
I woke up to this news delivered from Yann Lecun himself in the morning on facebook[1] and my gaped mouth can still be found for onlookers to witness I suppose!
LIKE THIS IS IT FOLKS!
Edit 2
All those back in my days muzzak folks lamenting about the quality of contemporary music can fuck right off because you clearly havent explored enough of the modern music landscape.
Dont you dare blaspheme saying modern music has stagnated or some drivel like that. It is outright offensive to folks who are pushing the boundaries like say for example The Ex from Netherlands https://www.facebook.com/theexband https://www.theex.nl/news.html
Just because you and the other soulless people you fraternize with are ignorant of all the innovative stuff thats going on, we have to suffer through your opinion on the state of pop culture?
I don't know what the point of machine generated music is. Just destroying one of the few remaining ways for people to make a living doing something creative, I guess.
The promise of automation was to have machines do the things we don't want to do, so humans could have more time to do things we enjoy.
Instead, we are automating the things humans enjoy, and still leaving humans to figure out how to feed, house and clothe ourselves through the sweat of our brow.
Bingo.
This is a fun toy, but in terms what it means, you may as well ask an AI to pray. It's completely hollow in terms of the actual experience.
This could make suitable filler for idle games, ads, aquariums, and elevators. Not much else. Perhaps at best, a producer could use this to fill in the instrumentation behind a singer, but I have a feeling it's not there yet.
> The promise of automation was to have machines do the things we don't want to do, so humans could have more time to do things we enjoy... Instead, we are automating the things humans enjoy.
Damn. Never looked at it that way. It's still enjoyable to do these things, but perhaps less lucrative. I don't know, do professional musicians like arranging elevator music? I'm strictly an amateur who has never made a dime performing, so I really don't know if that would be joyful, soul-crushing, or somewhere in between. I just know what it means to me, and like I said, you may as well ask the machine to pray for all I think this amounts to.
The generative process is based on a combination of learning and randomness. The random part doesn't mean anything, but it's clear that it is far from just random notes. Do you think human music always starts from a meaning? It's just lucky accidents that sound good. We even retrofit explanations post facto to our actions, we can certainly compose music first and assign a meaning later.
Around 150 years ago classical music had a big dilemma - should music be related to concrete things or abstract? Should we put a story to music? So everyone wanted to know "what was the program?" (program==original author's meaning) sometimes composers would just hide it in order to instigate people to use their imaginations. It didn't matter what meaning the author originally assigned to it, better to try to hear it with beginners ears.
One point is that music fans can now make their own music. I think it's great that people can express themselves and it's not limited to those who put in 10k+ hours to master a single instrument. More people creating is a good thing.
So there isn't going to be an increased level of profound self expressions because of this. Quite the opposite, more pure noise for the purpose of farming ad revenue.
What's worse, and an aspect many proponents of AI generations ignore, is that by ushering people into this specific channel of caring more about prompts than all else, we are doing a real disservice to potential people who could have become serious masters of their realm. After all, "why learn how that music program works when I can just generate it?"
Things will erode and decay, things will come into being, things will change. This flux is so constant that in truth there hardly are any things, just the changes; for as soon as you step in the river a second time, neither you nor the river are the same as you were. Epictetus, maybe? One of those guys.
Likewise, music is inherently fleeting, yet it still makes sense. You can't hold music, yet there's still a sense of it being a thing that exists. Yet when it stops, it still somehow hasn't ceased to exist. The act of musical performance, even at a basic level, especially with others, brings us one step closer to something fundamental about the universe than other forms of expression.
Like I said elsewhere, if you could ask the machine to pray or meditate, it wouldn't be fulfilling for anyone. It would be hollow.
Is it really an act of personal expression if you've narrowed the "vocabulary of creation" to the stable Diffusion equivalent of "hyper realistic, unreal engine, 8K, masterpiece, intricate details"?
At that point would not the act of creation feel rather hollow?
I can create something that sounds decent in software like Ableton Live or whatever, but I can't play anything on a piano or a guitar.
I would like to make music, video games and movies, too, and AI lets me do that. I don’t need millions of dollars or years of training to make something creative anymore.
You never did... You just needed to get creative.
In what sense is that still you doing the creating?
Speak for yourself. I like music if it sounds good, regardless of who made it.
Have you been to a farm before? Have you seen a textile factory? Have you seen a construction site? How could you, with a straight face, suggest we are not automating those things? There are vastly more people working on automation in those fields than are working on AI-generated music. Automation in agriculture, construction, and textiles are massive industries. There are a lot of people in the world working on a lot of things.
Food is a different problem. We have access to very cheap calories, but the overall quality of nutrition is way down in advanced economies, leading to an epidemic of obesity.
Textiles is pretty much a solved problem. We have so many clothes, we give them away en masse in developing countries. I think there are very few people in the world without access to adequate clothing, and if there are I suspect it's a distribution problem.
But why exactly should that happen? By which mechanism? Every single company automates in order to increase their monopolies and profit, to generate more shareholder value. There exists no other mechanism, so obviously we will never do anything other than that.
Exploring the latent space of human music. It's a cultural mirror.
You're totally okay not feeling this angst. But so are the folks who do.
As I've watched the evolution of music generation with LLMs I feel like I just keep hearing drivel at greater fidelity. If you like it then by all means listen to it, but this is average or below. In some ways I think I prefer the more chaotic less coherent predecessors. They're a bit more interesting to my ear.
And as other posters have said: that doesn't really sound like Aphex Twin to me at all.
On stuff like art it's hard to judge objectively, but in things like code it's much simpler. Don't get me wrong there are cases where I find generative AI useful - but the hype machine and the unedited whole solutions are just straight garbage.
It sounds like output not resembling what you requested, and you're celebrating because for some random reason this particular prompt didn't sound totally horrible today. But it isn't intentionally making music, and it isn't particularly interesting music either. It's basically baby's first drum machine sort of stuff.
PRECISELY! And I find that magical. Better prompt fidelity, model zoo etc will follow soon.
However, that's from a sampled drum beat. I generally agree though that this generated snippet doesn't remind me of Aphex Twin much at all.
There's already so much art being made. Why have so much joy in ignoring all of it and focusing on generative AI instead?
Furthermore, there's another work posted on that FB account that has the caption:
"While I'm concerned about the possible impact on society - especially on the jobs front, I cant help but grin as the edifice of human exceptionalism is shred apart with every passing day."
Is that you? Why the misanthropy? "edifice of human exceptionalism"? As it applies to... making music?
One answer would be to create music that shares its roots with music that the listener already knows. This music could be enjoyable, but you can't exactly sing along to a melody you're hearing for the first and last time, so it has more limited engagement potential. This is an approach to composition that you learn when you study chord progressions and other elements in music theory, and it's what I'm sensing when I listen to the MusicGen outputs.
To draw from greater cultural context, you can incorporate folk and popular melodies that are widely known. Musicians love this trick. "Immature artists copy, great artists steal." MusicGen seems capable of doing this, too.
To promote a novel melody as something that listeners deeply cherish, or to innovate at the level of the theory, the social context has to be built up around the content after it's generated. E.g., when introducing a new song on the radio, a common trick is to play it between songs that are already popular; building up co-occurrences with songs that already have cultural significance. My challenge to Meta would be: can you use your platform to transform some of the model's novel outputs into familiar popular music? It would be an important cultural milestone if an AI-generated melody became a familiar tune that would be played in the café, recognized, and enjoyed.
Think of your favourite TV show.
If when you first watched that show, you were told no one else had ever watched it, or ever would watch it, would your engagement with it be the same?
Part of our enjoyment of art is the shared cultural context. Maybe saying "art" here isn't the salient thing. Maybe it's our engagement with ideas.
I personally haven't even considered this concept as much in relation to music, because while I do love to deeply engage with music and the shared narrative behind it, both real and imagined, I also just like to put on music that sits in the background as a tool to drown out noise while I'm working, walking, etc.
New generative music benchmark - popularity.
If music is an artifact appreciated by listeners, then any metric apart from whether the music is listened to would be a proxy — though I can appreciate the perspective of creating for the artist’s own sake, without a need to share the creations.
Popularizing some of the model’s outputs would reveal their merit against human-produced music, by allowing them to succeed or fail in attaining the same quality of cultural significance.
Then again, if the artist here is the AI research team and the audience is AI enthusiasts, then the music has succeeded in being heard and attaining cultural significance. It has been remarked that Schoenberg’s music was more often defended than listened to — music written by a theorist for an audience of theorists. I am a true fan of Schoenberg’s, though I can hear that the example outputs of this work are music that is meant to be accessible to the everyday listener.
In other news, goodbye Youtube audio library, this is pretty good
At 3.3B parameters this should be running locally, right Meta? (Yes it does, instructions on github)
I 'm not sure I've seen any MIDI LLMs , wouldn't that be more fun to do ?
What worries me, is that a good enough model will take away the incentive to write music for many, and as a consequence it will also remove performers. This will reduce demand on music teaching and instruments, which will then both become nearly inaccessible. Since learning music isn't a question of following a few youtube videos, this will leave the world with just AI music.
Jazz and classical music are probably exempt, since it relies on subsidies, their audiences care about the actual performance, and AI compositions will not draw enough of a crowd to make it financially interesting.
But popular music will suffer, and that's what makes development of these models straight evil.
Then 2 months later, they weren’t.
You’ve got to start somewhere.
Then again, in this case I don't mind. I'm sure someone like Simon Posford could do some really wacky sampling based off of this.
Don't see myself using it to make music just for my own listening though(not much of a composer). That's still a long ways off.
2. It doesn't have an "interrogation" feature for now (unless I missed it), but I'm sure something like this is possible.
If you took a DAW project file of a song.wav that was completely written and produced digitally using virtual instruments and compiled all of the parameters a user had to set to achieve their output.wav into a .csv file, you may be surprised to see (1) how few parameters were used (2) how often those parameters are unchanged from their defaults and (3) the amount of those parameters that would be expected in any other project file.
When you break it down, you really only have 6 layers to parse, all of which are dynamic but within a relatively small and consistent sandbox, at least relative to image generation.
1. composition layer - the midi or notes of the song. 2. arrangement layer - the selection of instruments used in the song and the division of the song's midi to the song's respective instruments. 3. instrument layer - the parameters of each instrument, such as a synth path or a virtual piano's room setting. 4. post processing layer - the effects placed on the output of each instrument, such as reverb, compression, delay, ect. 5. mixing layer - the volume of each instrument + post processing channel 6. mastering layer - processing on the master track
All of these things are more or less standardized. Developers always add their own flair (read: custom parameters) for their plugins, but they can be decompiled to be a composition of each of these layer's fundamental parameters. All these parameters + the midi of a song would be a few kb
I feel like a LLM trained on the parameter sets which interacts with the software used to manipulate these layers could really produce amazing tools and open the door to writing high quality songs to everyone, just as other AI products have opened so many similar doors.
The DALL-E app for music, in my mind, probably wont be a text based description -> .wav output. Instead, it would be the generation of the elements of each layer with options that can be listened to in real time using whatever VSTs were used in training. When you ask ChatGPT to write a complex python script, it starts with an outline of all the methods in the script as placeholders and then takes you step by step until you're done, then you troubleshoot it or flesh it out. The best part of a generative music like this is that it leaves the user really only with having to decide if something sounds good or not.
As a mostly musically illiterate producer myself, I've produced hundreds of songs and a few albums without ever really learning how to do anything other than manipulate the parameters. When I started learning to produce music I was 15 years old and knew nothing about music production. But, I was really good at was using computers and software so I learned to play the DAW, the plugins, and the sample packs. The only layer that I couldn't learn through learning software was the composition, the writing of the midi. Fortunately, the midi of a song becomes very easy to brute force over time, so I learned to brute force midi. Once I became efficient with my workflow, producing music became a task of "make this idea sound good." And without ever really feeling like I was a musician or composer, this became an enormous passion and outlet for me that I did every day for a decade.
I was able to do this because at its core, all the mechanical parts of a song are simple machines and a song's quality is the way those machines are used together. As an outside, this feels like a workflow that would be very machine learning friendly. But I could be wrong!
Seems to be more useful as an "assistant" for music producers, similar to how Copilot operates.
Right now the generation still sounds a lot like loop packs smashed together anyone could technically make. But it is practical for anyone who really only cares for that style of sound but do not themselves have the familiarity to do it. Now they can just say what they want and hit regenerate, skipping the latent feedback cycle of iterating with humans or sifting through song snippets.
My opinion on this style of content is that ai generation is simply accelerating us to the inevitable end of generic digital content, it isn't really changing it. It just happens to be also the optimal interface for discovering and not just generation.
> In the U.S., Boléro remains under copyright until 1 January 2025
Producing this kind of anti-music is pretty soul-destroying for a musician, anyway, so the machines might as well do it. We can then spend our time working on stuff that means something, and if we’re lucky and it connects with enough people, make a living from it.
This space will belong to scrappy shadowy decentralised organisations who let you type "give me a filtered french disco song using mizell brothers era johnny hammond jazz funk samples, lil uzi rapping, with a thundercat bassline and crooning"
There's also a relative dearth of royalty free music for independent content creators to use. AI would enable them to produce better content on a limited budget.
People who enjoy creating music from scratch will be unaffected - recognition and financial rewards are tiny already for most.
https://soundcloud.com/obie for reference
I can imagine an VSTi that just takes a prompt and generates midi tracks. Something like this is surely coming in the next couple of years.
Since the advent of transformers, and this idea of using text models mapping the natural language space to tagged music samples, and the music tokenizer acting directly on sampled audio stream (the bits of a .wav file, essentially) all the cutting edge work is going that route. Because it is producing high quality, finished audio streams directly. And I think part of it is because there is way, way more training data for actual audio than there is for MIDI alone (there are tons of free midi sites out there but a lot of it is garbage and it pales in comparison to what is already sampled and tagged in real audio libraries).
I imagine what will happen is... within two to three years these LM-transformer-music models will get so good that the audio will be damn near spotless sounding, and there will be additional methods developed to synthesize with more control directly with the models, to the point where wanting MIDI so you can use your own HQ synth isn't needed, because if you want "the lead synth to sound less digital and more like a classic Minimoog Model D" you just add that to another 'Music2Music" pass and out pops your sound.
For those who still want MIDI there is still work being done on traditional audio-to-MIDI modeling and I think you'd wind up just using that in the chain.
It's an amateur radio thing btw...You go up a mountain and make use of the good propogation to call other hams. You accrue points that are as valuable as HN points are!
No such thing as chambers music or a music. Maybe "a sad piano and cello duet" or "sad piano and cello chamber music" would be a better prompt.
I sleep. And these are only 1-3.3B param models, that makes no sense.
"The weights in this repository are released under the CC-BY-NC 4.0 license as found in the LICENSE_weights file."
Combine it with AI voices and voice cloning and so begins the further devaluation of musicians and artists.
Might as well accelerate it and see what happens. What could possibly go wrong?
That being said - it's happening, and nothings going to stop it regardless of which side of the fence people sit on.
(I’m not saying there’s reason for AI development in general to stop, but these generative things that are designed to slot neatly into the role of human artists specifically have no reason to be developed further beyond proving it was possible, and that happened a while ago.)
I think you might enjoy Neil Postman's book "Technopoly", which discusses the subject of weighing the pros and cons of a subject instead of just diving in headfirst every time some new technology is developed. His YouTube talks are also great.
Well, the European Union is already working on a legal framework for AI. It happened with GDPR and it will happen again.
LOL, I love Americans and America but seriously? Like what is already there is not enough :D
Surely the musicians suing each other aren't the ones that are now planning on training an AI on other people's music?
haha, did you conceptualize music ex nihilo?
We’re autonomous entities capable of higher reasoning, limited in time, attention and talent and eventually die allowing new people time to flourish.
We also pay taxes and make silly arguments about software and humans being no different from each other because it justifies our ability to play with cool toys without considering the impact on other people. Corporations aren’t people though because they aren’t cool like AI.
A profitable startup? That might be asking too much though
Maybe someone will say this will let people who aren’t musicians express themselves by creating music. That’s not true. It’s as true as hiring a musician to make a song for you, given a description. And nobody would say that the person who hired the musician was expressing themself.
What happened to automatic the boring things?
Instead they seem to be all in on washing out any hope in creativity and pointing people to put all their hope in minting and munging “code”.
It’s so myopic and short sighted it hurts my soul. I don’t understand at all. All that money, all that knowledge and talent… and this and stupid headsets strapped to peoples faces is the game? God dammit.
I'd say it lets me, a non-musician, create generic tracks for other projects that need music, but don't need the Lord of the Rings soundtrack.
If you want the insane soundtrack, you still need an artist, as this project demonstrates.
That'd be pretty cool for example.
This kind of automated "filler" music has been around for decades, and is usually used for exactly that - filler. It's pretty much the stock photos of music.
And that could be a good thing - suddenly content-creators don't have to spend money or energy on purchasing that kind of stuff.
If you've ever seen youtube automation videos - typically those "TOP N" list vids, they always contain some kind of muzak-style soundtracks.
Unlike classic "hiring a musician", here it's practical to "hire" the (robot) musician 10000 times with a feedback loop between the model and the prompt writer, iterating and picking the best output(s)... which looks like a similar process to other exercises considered art forms.
Man wants to dominate nature and always has. I don't think this is particularly difficult motivation to understand, as it seems omnipresent.
And every tech advance has side effects , usually in unforeseen ways
Something tells me AI isn't going to rescue us either. I just sampled a bunch of these generated tracks and they immediately remind me of the average, mediocre, soul-lacking content that most music and film is today.
When I was 20 I was a music snob into Aphex Twin and weird IDM. I thought all pop at the time was crap, like you seem to. But then I heard, I mean like really heard, "Bye Bye Bye" by *NSYNC and seriously that is a good song!
I'm 40 now and I think it got way better even since then. Pop is so varied now! I really don't think music as quirky and weird as, say, Billie Eilish would've made it to the top of the charts in the 90s. I'd say that music like hers (and many charting artists of her generation) is a testament to how broad and compelling pop music has become.
My generation thought their parents' music was shit, my parents' generation thought their parents' music was shit, and so on, all the way until at least the invention of Jazz. But the average Gen-Z'er thinks all the music is great! They invent new genres for every song, they wear Metallica t-shirts in 2023, and they mix 80s disco with 00's Brit rock like it's just what people do.
And don't forget there's an endless long tail of music out there. There are so many good musicians and plenty of them have a sufficiently fancy label deal to be on Spotify and the likes. And otherwise they're still on Soundcloud, Bandcamp and YouTube. It's worth a deep dive!
If this appeals to you, it's worth checking out Japanese music from the Showa era to present. They've long mixed styles in a way other music markets have not. You can hear city pop songs from the 80s with metal guitar solos, jazz progressions a samba beat and synths, all in the same song.
However, I think modern rec algorithms (like the Netflix home page) are recommending more mediocre stuff than the old system, and the streaming boom did produce an abnormal glut of junk.
Anyway I think AI is going to spawn a music remixing/game modding/tv extending renaissance. They perform much better when pointing them at a good source (as you can see with the melody conditioning samples, and other stuff like sd img2img and finetuned llms).
But for movies and TV? Where do I find the good stuff? It seems Hollywood is creatively bankrupt and just milking people off boring franchises and cheap nostalgia through crappy remakes and sequels. My eyes rolled to the back of my head when I saw an ad for a show called “how I met your father” on Hulu.
There’s a trove of incredible foreign movies and TV shows out there. Scandinavian and Asian (Korean in particular) content has a really good hit to miss ratio for me.
For examples, check out international film festival nominations and winners.
Reelgood is a good one, sort by IMDB score (which is somehow still kinda working as a metric) or the reelgood score which is a popularity among enthusiasts kinda ranking. You will find tvs gems streaming services criminally and inexplicably never recommend.
But "old school" recommendations from TV /movie buffs (like the tvtropes community or various forums) are still a good source.
As far as content goes, there has been a ton of excellent stuff just this year across movies, TV, and anime. One “organic” way to start is to look for recent recommendation threads on Reddit for a movie or show you really like.
Mainstream entertainment has always converged to mediocrity.
Automation tools and fine grained computed metrics have rounded off the edges of emotional experiences. See Bobby Kotick about taking the fun out of games: https://www.escapistmagazine.com/bobby-kotick-wants-to-take-...
Him saying that is around the time the music compositions start becoming similar. The mentality was not constrained to games.
Nothing is allowed to be it’s own thing anymore. It has to be hypernormalized to have enough reach a billionaire CEO can profit from.
MBA-ification of reality.
> The act of shifting a song’s key up either a half step or a whole step (i.e. one or two notes on the keyboard) near the end of the song, was the most popular key change for decades. In fact, 52 percent of key changes found in number one hits between 1958 and 1990 employ this change. You can hear it on “My Girl,” “I Wanna Dance With Somebody,” and “Livin’ on a Prayer,” among many others.
To me, this just reflects one set of songwriters' cliches being replaced by another. Not necessarily better or worse.
Music especially just feels flat. Maybe that's just the style now, and I'm old and can't appreciate it.
Honestly, gaming is in a similar rut although not quite as bad thanks to VR.
The music is fantastic if you just look a little.
There's so much great new music being made every year, new genres and ideas, etc. Film music seems better than ever recently. Especially for TV series. Lots of new styles emerging there too, see Mac Quayle for instance.
The really good, modern music was almost always on the fringes, and there's more of it now than ever before.
There might also be more garbage, but there's no need to listen to it.
There is an overwhelming amount of good music out there. Pick an album top 50 list from 2022, for example fantano's, or pitchfork, check out bandcamp's staff picks, listen to other musicians that are on the same label as your favourite band, keep an eye on things like NPR Tiny Desk, KEXP, la blogothèque on YouTube.
Just start listening. You are almost guaranteed to stumble upon something you like. It won't come to you algorithmically but the effort required is really low.
My favourite new album I discovered last year was Immanuel Wilkin's The 7th Hand [1], I stumbled upon it by going through a top 20 jazz albums of 2022 list to see if I had missed anything, and it immediately jumped out at me as being exactly the shit I'm into.
https://en.wikipedia.org/wiki/List_of_Billboard_Hot_100_char...
ETA: And "worse" in these studies tends to be defined in terms of measurable qualities where contemporary pop music most differs from "classical" music.
Computers can already generate every possible waveform that can be heard by human ears. There's nowhere else to go. Anything above 16kHz is useless and can't be heard anyway.
The brothers Rosling have a nice talk about how all the big stats are improving globally (gender equality, education, health, extreme poverty, life expectation)
I agree with you. When gwern investigated AI folk music in 2019, I realized it could generate a wonderful variety of music, full of soul. Be sure to listen to several tracks before making up your mind. My favorite is “crossing the channel”, since I think GPT made a mistake at the beginning, and then generated the most reasonable sounding not-mistake, which turned out to sound so cool.
My goal was strong, memorable melodies. Star Wars, not Marvel. GPT can come surprisingly close, if the input data format is right. Unfortunately I don’t think anyone except gwern has noticed that the input format is crucial: https://gwern.net/gpt-2-music
To be immodest for a moment, my work serves as an example that it’s possible to do it, and better than anyone else, long before they figure out how. Many examples of this pop up throughout history, and I am gratified to be a small but real one.
One day he posted something that sounded pretty amazing, and I was blown away. “More like that, please.” It had chords in it.
He didn’t pursue it past that. I did. So it’s possible that no one is aware of how crucial the input format actually is to the success of the music that I was able to produce.
(And “produce” is a fair description here — choosing the instruments was really important, and the model didn’t do it. It wasn’t as easy as press a button. It felt like I was suddenly a 15x music producer, since I made all those tracks in one night. Such is the power of ML.)
It has been easier than ever for any individual to create content of any type, and there will always be gems.
AI isn't meant to "rescue" you from this problem, at least not in the present stage. You are looking for a mission that was never claimed.
This is a great song for contemplating loss of a loved one.
I'm not sure AI music will ever reach these heights because it will have trouble understanding death.
This model is just the equivalent of GPT-2 for music. It's not the GPT-4 yet. Music is trailing a few years from language. Used to be that language was about 5 years behind vision. Now language is the top.
Look at the samples used by Daft Punk on their Interstellar 5555 album. Plenty of originality there.