MusicLM: Generating music from text
arxiv.org
arxiv.org
They really need to bump up to 48 KHz so all the music doesn't sound like it's being played over a telephone. A factor of two in the cost shouldn't be prohibitive. So much of the audio generation stuff I've seen has fatal flaws like this baked into the dataset and/or training process that ensure the output can't sound good even in theory, it's kinda frustrating.
It's also frustrating that nobody AFAIK has trained a music model on an actually large dataset. We're training large language models on a significant fraction of the text of the whole internet! Where are the audio models trained on a significant fraction of all recorded music? This one was trained (in part) on a dataset of 280k hours of music, if I read the paper correctly. I don't think that comes anywhere close to being a significant fraction of all recorded music.
https://ooo.ghostbows.ooo/about/
Music is on all the usual streaming platforms, the album is called Shadow Planet, band is The Cotton Modules. Also avail on their website:
There is a million Industrial techno sounds, repetitive, hypnotic rhythms IDM tracks that absolutely no one listens to anymore. Most the founders of the whole genre have moved on from lack of interest.
We have been able to do really good generative music in Reaktor for the last 20 years. Much better than anything on this page but no one cares.
People in general want to hear the same slight variation in music over and over as background noise.
The people that are saying how great all this is will be bored with it in 2 weeks or less. Pure technological kitsch.
These models are still in their early phases - If AI was combustion engines, we wouldn't even be to railroads yet. Despite that, they're still fun to play with and useful when used correctly.
I was born too soon to explore the stars, born too late to be a 17th century pirate... but born just in time to explore these incredible neural networks come to life.
https://news.ycombinator.com/item?id=34541693
Perhaps they could be merged?
I think this showed up on my recommendation. I didn't know anything about tracker music, but this video explains it chronologically and very well.
Do we still need humans to knit and handwash our clothes? To live in a caboose or lighthouse? To operate our elevators? To deliver the milk?
Why should anyone have to learn to draw in the future when machines promise to do a better job? There are so many better uses for our time, and too many things to do for our short lives to handle.
I want to compose music, despite no practice or formal training. I want to make an entire movie by myself with no other humans involved. Tech that enables these things will be empowering.
An analogy: chess AI's are clearly superior to human chess by any reasonable measure. I still enjoy playing with other people (even online, even anonymously) infinitely more than playing with an AI, no matter how well calibrated/tuned to my level.
I hear a lot of romantic notions yet I don't see a lot of people flocking to solo acoustic singer-songwriters!
An AI without new inputs to consume will not evolve creatively, and you will get bored of its output quite quickly, I think.
Yes, today ML can write you some run of the mill music and design assets, tell you how to develop the game, it can even play your game and write a review on it, but what’s the fun and/or point in that?
especially the conditioning on humming and whistling examples are cool, but to bad they use very common melodies for that so it's easier job for the model and harder for us to judge how well would it work on less common melodies.
100%. Now if you add the detailed "Painting Caption Conditioning" to the mix, you can create a melody in the style of… an image, which, in a way, is kind of a controlled artificial synaesthesia [1].
I can’t wait until they get to the point where they’re more composable or auto-accompany given an acoustic guitar and vocal input.
I find it fascinating that MidJourney can make a 3D model of my face from a low quality image, rotate it in space, apply it on someone's else body, add coherent shadows and backgrounds, with a very credible result, and yet an AI cannot generate a decent song, which is 1-dimensional and has probably much less internal modelling to care about?
One reason I can think of is because eyes are "integrators" and ears are "derivators". That is, that human ear is very sensitive so small differences, whereas vision cares more about the ensemble? I don't know, but I think that AI music will come one day. It may not be as great as human music, but it will suffice for, say, putting a music background for your startup cheap marketing ad.
Writing somewhat novel music to a formula arguably has a much lower skill bar than producing good representative art (especially if you allow sequencers as composition and disallow basic digital retouching of photos as visual art), but producing something that genuinely stands out may be harder
There is an ocean between “I like that sound” and a final, produced piece of recorded music. Much of that ocean being ineffable.
Trying to reduce it to a set of parameters around the final waveform and you’ve missed the entire point.
> but it will suffice for, say, putting a music background for your startup cheap marketing ad.
We need less of that—not more. It’s like a climate disaster of the soul.
The only thing it will really F is that bridge that art built to bring people in closer touch with themselves.
Thinking we can shortcut self-discovery, individuation and essential human connection with an algorithm is insane hubris. Icarus’ wings are melting.
Just like a 3D model is really only 2D on a screen, music can encompass so much more than what is heard on the surface level with no imagination.
For music, I think it's partly an academic question of "can we do it" rather than trying to maximize immediate practical usefulness. There's already quite a bit of work on symbolic music generation (mostly MIDI), a lot of it quite competent, especially in more constrained domains like NES chiptunes or classical piano, so a full text-to-audio pipeline probably seemed a more interesting research problem.
And for a lot of use cases, where people might truly not care too much about tweaking the output to their liking, the generated audio might be good enough; the examples were pretty plausible to my ear, if somewhat lo-fi sounding (probably because it's operating at 24kHz, compared to the more standard 44-48kHz).
In the future a more hybrid approach probably makes sense for at least some applications, where MIDI is generated along with some way of specifying the timbre for each instrument (hopefully something better than general MIDI, though even that would be fun; not sure if it's been done). I'm sure that in the near we'll see a lot more work in the DAW and plugin space to have these kind of things built-in, but in a way that they can be edited by the user.
Other libraries of (Royalty Free, Public Domain) sheet music:
Explainable artificial intelligence: https://en.wikipedia.org/wiki/Explainable_artificial_intelli...
FWIU, Current LLMs can't yet do explainable AI well enough to satisfy the optional Attribution clause of e.g. Creative Commons licenses?
"Sufficiently Transformative" is the current general copyright burden according to precedent; Transformative use and fair use: https://en.wikipedia.org/wiki/Transformative_use
That being said I have seen very little in the AI transcription space (Listening to music, and outputting legible sheet music while identifying instruments) especially when it comes to multi-instrument music. I've also seen very little in the midi generation space aside from what you mentioned.
It would be nice to see midi to midi AI remixes, analogous to how Chat GPT can (sort of) turn wikipedia into rap lyrics.
Though something end-to-end would probably make more sense; text is a lossy way to encode speech, as it doesn't capture a lot of nuance that goes into spoken words. Interestingly, the same could be said of sheet music and even MIDI in the music realm.
As for transcription for multi-instrument music, there's been a lot of work on that front, and a lot of progress that's already finding its way into commercial products. The latest I've seen in this domain is https://samplab.com/, which has some compelling videos of not only transcribing multi-instrument audio but editing it on a note-by-note basis.
If the argument is that "maximize immediate practical usefulness" is not something interesting, I would say it makes sense to say "I don't understand why this approach is pushed". Anyway.
In music, 'fixing' problems in sound is way more costly, so a producer would say that it's not something you can work with in any reasonable manner.
Pretty sure the Beatles never handed George Martin any midi files. What’s the symbolic representation that captures the tone of every bend in a Hendrix solo? Did Daft Punk go back and grab the raw master stems of the old vinyl recordings they used to assemble their tracks?
Music producers have been astonishingly creative given inputs in a vast range of formats. Sheet music and midi are one tool, but ultimately it’s about combining sounds in the mix isn’t it?
The issue with combining sounds in the mix is the issues add up in a super-linear way. Because of that you could even say that modern music producers tend to care too much about perfect sounds.
The Beatles did not have the same production requirement modern music have since there is high-fidelity equipment. The music industry as a whole changed a lot. Modern music production have multiple passes just to clean up one "perfect vocal take", like de-essing, breath removal, etc. This is insanely more advanced than what was the standard in the 60s. I'm not saying it's better either.
>What’s the symbolic representation that captures the tone of every bend in a Hendrix solo?
That's not how music works. You refer to the score, and the precise way you bend your tones is up to your interpretation. Even if you wanted to copy one-to-one Hendrix, you'd copy the style, not one performance exactly, so capturing every little bend does not really matter.
Which could be an AI project: Hendrixify a guitar track.
My point is that performances are as valid an input to music production as anything else.
And that these generative snippets are potentially a source of ‘performances’ that producers will tap in creative and interesting ways.
Yes! That's what I figured when I started my project to create an AI assistant for melodies (http://melodies.ai/). I'm quite sure that splitting it up is the way to go, and that the first AI hit song will not be created as a full song but by combining instruments/vocals.
I want someone to train it on public domain music. Kind of like the YouTube Audio Library but I assume that's not exactly the right license for this. But with sufficient effort someone could make a lot of recordings of public domain music for this purpose and build something that the RIAA thugs couldn't actually touch.
btw, some of you might find interesting, diffusion-based / generative concept album: https://wielu.bandcamp.com/album/vectorstep-ep
Also maybe a silly question, but what's the legal ramifications of downloading these YouTube videos and training on them yourself? Google must have some rights, but what about people outside of Google?
Custom generation of samples from text alone seems revolutionary.
what a joke
"""We acknowledge the risk of potential misappropriation of creative content associated to the use-case. In accordance with responsible model development practices, we conducted a thorough study of memorization, adapting and extending a methodology used in the context of text-based LLMs, focusing on the semantic modeling stage. We found that only a tiny fraction of examples was memorized exactly, while for 1% of the examples we could identify an approximate match. We strongly emphasize the need for more future work in tackling these risks associated to music generation — we have no plans to release models at this point."""
Also, is this technically a vocoloid? :)
Anyone want to share their experiments with MusicLM, feel free to join community-fan discord and subreddit:
BTW, I really hope this gets integrated in software like Ableton similarly how image generators get slowly integrated into Adobe Creative Suite.
How can we experiment with MusicLM?
The paper’s last sentence says, “We have no plans to release models at this point.”
But, I guess, one way to make money currently.
I don't know about this. I know professional musicians and they say that they make little mistakes all the time, but part of being a professional is being able to pretend those mistakes didn't exist because the audience doesn't know what it's supposed to sound like.
Very true, however it's something musicians have to specifically train for, because to a trained ear they can become painfully obvious and it's tempting for the musician to (at least) cringe in a way that the audience will detect. Musicians in the audience, though, can often tell when you flub a note, regardless of how you play it off. That's really the bar we are looking to meet.