Turn ideas into music with MusicLM
blog.google
blog.google
https://share.getcloudapp.com/7KuzO6QO
Let's make some tropical house, use the lyrics "call my private number" as the hook, the vibes should be disco, light, fun:
https://share.getcloudapp.com/2Nub4m6B
Let's take some classical music and turn it into techno:
https://share.getcloudapp.com/NQuW6lBP
tropical house, 120bpm, bongos, wind chime, flutes:
https://share.getcloudapp.com/2Nub4mDq
Last one is pretty cool ^^
Well I'm just going to go ahead and pretend I didn't hear that.
According to the paper[0] "We found that only a tiny fraction of examples was memorized exactly, while for 1% of the examples we could identify an approximate match." FWIW I'm pretty sure I heard a fragment of Kalinka (a folk tune in the public domain) in one of the samples I generated.
They all suck. Nowhere near Jukebox. Depressing that seemingly zero progress has been made since that.
Remember all the people who said, in January and February of 2020, this "novel coronavirus" is not a big deal? Because it affected only a tiny amount of people? Because "much more people die from the flu"? The fallacy is that they ignored the speed at which it progressed.
From a music producer standpoint if that's the quality for classical music, could definitely create some samples.
It has multiple structural layers of a sort, which is progress over a few years ago. But they're still a long way short of the huge but intricate structures in real classical music.
It seems like it's trained on entire soundwaves. I'm curious if you'd get a better result by training it on transcribed MIDI and then taking the output MIDI and plugging it into VST's.
Seems like you would still get that "central brain" compositional approach without the garbled sound quality and unidentifiable instrument noises.
The information density of music is much higher than that of text or still images. So something like this is still tech overreach.
It's not unlike how the visual AI can do 'Greg Rutkowski', but has a hell of a time being an actual concept artist in a functional way. If the cliche soup is well defined, you're pretty much all set, particularly if it's not a genre that requires a lot of character.
That depends on how you encode it. As a sound file (.wav, .mp3 or something like that) it's hard to compress but as for instance a midi file it can be very compact. Music is hard to make and hard to reverse but it is relatively compact in terms of source material if it can be expressed as midi.
sexy sax - Oops, can't generate music for that
drill rap with gunshot - Oops, can't generate music for that
song with farts - worked but it contained no farts sadly, just some bad techo
sad song with crying - worked but contained no crying, just screechy violin folk music
Jokes aside, it’s still very early days for this kind of generative AI. I see a real use case for virtual band mates that play along with whatever type of music you play when jamming at home, for example.
I fear the path into "low/no effort" versions of music scores, art, etc is a dilution of great works that actually makes a worse product than if you went to the trouble of finding someone to do it properly. It leads to shovelware and I don't think we need more shovelware. We need more high quality high intention music and games and images, and personally I think thats going to come from gifted individuals who put in the effort to learn it, not a neural network that can't experience what you want the music to make you feel.
If I logged into a video game and I got the music in the above post it would detract from the game. You'd be better off finding open license or cheap music that a human has made. its way better, and would reflect better on your game and design process.
> The use-case, I think, is more along the lines of generating adequate background music for a my-first-self-published-videogame project with next to zero effort
next to zero effort is the goal. shovelware.
Having worked on a few games doing music, the devs were super passionate about a quality product, frequently over-engineering, and when then-available bog-standard audio middleware didn't do the procedural mixing we wanted, wrote a custom system just to get it spot-on.
I also know a few major AAA legends who went indie who won't publish shite.
Equally, if the original commenter is happy with x being good enough then that's valid. Maybe the game isn't for you?
Keep in mind my initial reply was directly replying to the "What happens if you've built the next hollow knight" in your post. The implication of my first reply is if you're at such a level you can do that, you very likely aren't going to settle for your music letting you down, which is arguably the other 50% of what a modern game is against the visuals. You're focusing on the commenter being satisfied by AI music. If the game is that good, you'll get a budget to have the music not suck, whether that's via XGP paying to finish the development, a revshare with a composer, a grant, or even a publisher.
The tl;dr is if buddy thinks his game is anything more than a learning exercise or something he simply enjoyed making and has potential to actually be a great piece of art or a decently selling product, he absolutely does not have to settle for mediocrity even without a budget.
Clearly I'm not - as I mention hollow knight.
>Equally, if the original commenter is happy with x being good enough then that's valid. Maybe the game isn't for you?
Clearly it isn't for me, as I'd have to listen to the above sort of garbage music and would hate the time spent, regardless of the game itself.
> The implication of my first reply is if you're at such a level you can do that
There are lots of people who have great ideas and even iterations on a game that isn't there yet, but who fail in graphics, audio, and overall design because they underestimate how important those are to user experience.
I would never say "I'm so glad I have access to copilot, now I can make a game with next to zero effort" but the guy I replied to thinks he can get a score worthy of being in a game from a machine learning model.
The easy part overrides everything and we get (or are going to get) huge collections of shovelware from people who see how easy it is to produce.
So this is not just empty PR by Google.
Its part of AI test kitchen, if you were already part of that for playing with BARD and stuff you just need to log in and give it a go.
If not, you'll have to get on the waitlist. When I asked I got access within a day, but might be much slower now with so much interest.
Here are a bunch of examples of the music: https://google-research.github.io/seanet/musiclm/examples/
In my opinion it isn't exactly great. It doesn't do much creativity. Eg. I asked it to make a dance song from a club but have classical instruments as well. It was unable to do that. It's mono and can often be out of tune.
We will see how it improves, but this certainly isn't taking away jobs for musicians in it's current state.
Part of the problem is the awful sound, would be much more interesting to have scores generated by AI and then have musicians record it. As it is you can't really tell why it sounds so awful.
I.e., I want to define the 'problem'. In musical terms: I have an idea, it's 8 bars, maybe 12 or 16. I don't like what's happening at bar x. Give me 3/5/7 options for something different, but maintaining the instrumentation/orchestration I've established. Give me some "temperature" controls to depart, or not, from what I've composed thus far. Etc, etc.
It will be one of the more funny and novel way to lossy encoding music.
And I just checked, GPT4 will "make music" in ABC notation, nothing like "Tropical House 120BPM" tho. It is pretty lousy artistically, that it can do it at all is absolutely amazing.
Here is a remix of Cooley's https://web.archive.org/web/20230513053704/http://pastie.org... which is the only ABC notation song that doesn't sound like ____ https://abc.rectanglered.com/
GPT4's remix has me in stitches!
Because there's no relation between the note you play and the chord mapped to it, you can just noodle randomly without thinking of what you're pressing and just use your ears.
FWIW, the various diffusion models (tend to) use cross attention to an attention-based approach.
Algorithmic music composition has usually been split into two:
1. Generate notes (re: theory, genre)
2. Generate sound
(i.e., EMI[0], Kulitta[1], MusicNet[2])
Now we are doing both at the same time, and backwards. The model isn't (necessarily) going "write melody, then generate the sound", but rather, "here are 500 songs that are described with X, 500 with Y, and you want XY, so we'll combine these two" :)
(This is my best understanding, so feel free to correct)
[0]: http://artsites.ucsc.edu/faculty/cope/experiments.htm
The problem is that coherent musical structures are much more constrained. You can't just XYZ... into a space and get something that makes sense.
That will kind of work for low-density music, which includes a lot of landfill dance + subgenres. But these statistical models are blind to larger and more complex structures, and completely unaware of cultural context and semantics.
It's actually a harder problem than language modelling because the spaces and the grammars are much larger, especially once you start including sound quality and production values as well as arrangement and core composition.
We are already drowning in music, you can turn on Spotify and have enough music to fill a lifetime. Yet new music is still being produced, why? Because music is ultimately a psychological experience, the human connection is a not-insubstantial part of the experience.
There’s a place for AI in music but it has to be white box, there needs to be scope for a human to jump in there, modify things, and make it their own. Otherwise, who will care?
And an infinite stream of music is exactly what I want. I don't want to curate or search. I want to feel and ask and get.
(╯°□°)╯︵ ┻━┻
* It really doesn't understand what makes good fiddle tunes. So bad.
* It's somewhat better with electronic music, but while it's locally solid (each ~2sec bit is good) it's jerky when you listen for a longer period.
Back when I used to play in bands, I’d often jam with a mate who was a drummer, he’d wait for me to start with a riff and then join in.
Sure, I can record some riffs at home and add realistic drums myself in a DAW, or use a basic drum machine, but the magic was in the unique unpredictability of how he’d interpret my riff and the beats he’d come up with on the fly.
As I got older, I stopped playing in bands, lost touch with this particular mate (who also stopped drumming) and generally don’t play guitar as much as I used to.
Give me a generative AI bandmate that listens to me while I play for a second and joins in and you have my money.
I believe the (tech using) public has had enough exposure to AI generated voices in music that this experiment will likely seem disappointing.
Perhaps I was over-cautious or unfair in this case (I'm all in on MS/OpenAI stuff) just a gentle suggestion to give it some consideration if you don't normally pay it much mind.
I think that something like this could be awesome given the right training data and model but this isn't quite there yet.
I could see this being used as a tool for song writers as it is currently it could be useful. Tell the AI a prompt listen to it and get some ideas. I do not think it is coherent enough to write a full song but I downloaded a few prompts and thought there is something good in this snippet.
This is good work whoever did it.
https://soundcloud.com/memorecks/sets/the-birth-of-aifi
Falls apart at times though a big step above riffusion ibiza 3am late night lofi which wowed the world 5 months ago.
Told it "bass heavy dubstep to rock out to at a festival late at night" and got this: https://share.getcloudapp.com/L1uDmyGB
I'm impressed! :D
Piano worked fine for me too. I haven't spent too long with it yet but I'm sure it can be done, even if not consistently.
if i had this i would let people make music with it for free, put it under a de-facto license that allows google & the creator to distribute it, host it on a semi-social platform, and take a cut of the money that rolls in
but that's crazy, its not like they already have a massive video platform to host music videos
Add in the fact that I was denied access to the Bard trial due to it not being available in my country and I'm pretty uninterested in Google's releases in general these days. I had the same experience when the new Pixel Fold was released and instead of getting to see the product page I was redirected to the default Google store home page for my country. I couldn't even learn about the product.
EDIT: I like MusicLM, I was trying it earlier today.
Google has tiers of support for products (i.e. "How comfortable is SRE that they will be able to keep this thing in the air if every SWE that built it gets isekai'd into a magical universe that isn't ours tomorrow?").
Anything they put out in front of Workspaces customers is max-tier support because their data shows that customers won't stop paying because they haven't gotten a new feature yet, but will stop paying if they get a new feature and it's broken.
I did not know this word, so...
IF you want “Google’s free consumer offering, but with more support and storage”, that’s Google One.
Workspace is a business-focused offering that favors stability, enterprise-oriented administration, etc.; it gets business-focused features that aren’t in the consumer offering, and it often doesn’t get new consumer features until they have stabilized and additional work has gone into enterprise administration for them.
Before Google One, buying Workspace (or whatever its name was then, its been through many name changes) as a way to get as close as possible to “Google consumer services (Docs, Sheets, Gmail, etc.) but with paid support”, but that was never what the offering was about, and there now is a clear and explicit offering that is exactly that, so continuing to complain that Workspace isn’t what it is explicitly not intended to be and what Google One explicitly is somewhat pointless.
No, you pay Google for a business-focused offering that favors stability and proper integration of management tools for enterprise, so you get stability and products that have had proper integration of management tools for enterprise.
You can pay (for more storage, support, and other benefits) for the Google consumer offering in the form of Google One, in which case you get the consumer offerings even when they are unstable and haven’t been integrated with enterprise management, but don’t get the business-focused components that come with Google Workspace business offerings.