Music Generation AI Models
maximepeabody.com
maximepeabody.com
> Vocal Synthesis: This allows one to generate new audio that sounds like someone singing. One can write lyrics, as well as melody, and have the AI generate an audio that can match it. You could even specify how you want the voice to sound like. Google has also presented models capable of vocal synthesis, such as googlesingsong.
Google's singsong paper does the exact opposite. Given human vocals, it produces an musical accompaniment.
In 2019, I built this thing called RaveForce [github.com/chaosprint/RaveForce]. It was a fun project.
Back then, GANsynth was a big deal, looked amazing. But the sound quality… felt a bit lossy, you know? And MIDI generation, well, didn't really feel like "music generation" to me.
Now, I'm thinking about these things differently. Maybe the sound quality thing is like MP3 at first, then it becomes "good enough" – like a "retina moment" for audio? Diffusion models seem to be pushing this idea too. And MIDI, if used the right way, could be a really powerful tool.
Vocals synthesis and conversion are super cool. Feels like plugins, but next level. Really useful.
But what I really want to see is AI understanding music from the ground up. Like, a robot learning how synth parameters work. Then we can do 8bit music like the DRL breakthrough. Not just training on tons of copyrighted music, making variations, and selling it, which is very cheap.
IMO this would be much more useful.
https://openai.com/index/musenet
Also, Synfire is a somewhat difficult to grok DAW designed around algorithmically generating midi motif as building blocks for longer pieces.
https://www.youtube.com/watch?v=OrtJjEiWBtI
It's not particularly well-known but it's been around for many years.
I'd link to some specific examples (easy to Google or search on GitHub) but I can't recall which models were more successful than others.
Having taken a class on Bach style composition in college - I think a rules engine with a random seed would certainly be much more successful at generating Bach style compositions than any neural network-based model ever will be.
And here is an interesting patent that Sid Meier and Jeff Briggs filed for their work on C.P.U. Bach: System for real-time music composition and synthesis https://patents.google.com/patent/US5496962A/en
I'll leave the ROM search up to whoever is interested :)
But there are lots of applications for music which parallel the applications of ai generated images - things that are more commercial in nature. The media is functional, for use cases such as commercials, or social media type videos, where people just need something for the ambiance and don't want to deal with copyright or anything like that.
Furthermore, the sound itself is crucial, so perfect calibration of a perfect sound is definitely a part of what can be clearly be sought (when you do not want to leave that to a secondary human process in the workflow).
The problem with LLMs for music (as currently implemented, not inherently) is people keep training them on complete tracks. They're very obviously being trained by people who are not musicians.
> Who are we to discern what is or isn't music?
Hopefully, people with good judgement, potentially capable of evaluating products.
The poster is clearly meaning "good music".
> Do you have the same opinion of text or code generated by or with the assistance of
There you go: the same way we note that some NN generated text is missing crucial qualities (e.g. intelligence), or that some NN generated images are missing crucial qualities (e.g. direction), you can surely admit the possibility that some NN generated sound may be missing relevant crucial qualities to the vetting of a good critic.
Well no, Feyerabend let himself be called an "anarchist" but clearly there is a "more scientific" and "less scientific" - they cannot give you a lecturing appointment at the LSE or elsewhere to just shrug.
> the listener ... as justification
As justification to what? A producer makes products for different markets: people may sell bars of sugar with appetizers and synthetic flavours, that does not make the product remotely similar to healthy food.
If you're generating the entire thing at once rather than stems or note data, you just have an elevator music generator which inexorably tends toward the lowest common denominator.
No one argued that one isn't of higher or lower quality. They're both music, as is evident by your choice of words. Processed foods are foods, not great for you, but they're still foods.
> The [original] poster is clearly meaning "good music".
All that really matters is whether users like what the generator generates
[0] https://suno.com/song/0caf26e0-073e-4480-91c4-71ae79ec0497
Fundamentally, a song can be represented as a 2d image without any loss
- the song file stored in binary, printed out line by line
- the sheet music for the song, ie instructions for recreating it
In AI/ML world we're usually thinking about encoding into a series of high dimensional vectors, not sure off the bat how to represent that as a 2d image
> Stem Splitting: This allows one to take an existing song, and split the audio into distinct tracks, such as vocals, guitar, drums and bass. Demucs by Meta is an AI model for stem splitting.
+1 for Demucs (free and open source).
Our band went back and used Demucs-GUI on a bunch of our really old pre-DAW stuff - all we had was the final WAVs and it did a really good job splitting out drums, piano, bass, vocals, etc. with the htdemucs_6s model. There was some slight bleed between some of the stems but other than that it was seamless.
My primary use is for creating backing tracks I can play piano / keyboard along with (just for fun in my home). Most of the time I'll just use the 4s model and will keep drums, bass and vocals.
Piano (or various keys), organ and some guitars (with effects) have a lot of frequency overlap. The model struggles there.
If this happens, main character syndrome may get a bit worse :)
AI models are tools, and engineers and artists should use them to do more per unit time.
Text prompted final results are lame and boring, but complex workflows orchestrated by domain practitioners are incredible.
We're entering an era where small teams will have big reach. Small studio movies will rival Pixar, electronic musicians will be able to conquer any genre, and indie game studios will take on AAA game releases.
The problem will be discovery. There will be a long tail of content that caters to diverse audiences, but not everyone will make it.
If you think Pixar is Pixar solely because they have an in-house software stack, you're missing the forest for a small shrub.
Good writing and good directing don't need hundreds of millions of dollars.
That's what costs millions of dollars.
Yes, they have an insane technology behind, but that's not what enables what they do. Humans enable it. Without human touch, that technology is just a glorified tech demo.
We're still keen to underestimate what an human adds to the process. We became insane in the pursuit of efficiency.
There are so many creators putting in intense work, and doing it on low budgets. You can't claim these folks don't have attention to detail. Check out A24, low and mid and low budget films, or independent films and you'll see a wide assortment of highly meticulous storytellers.
Pixar, on the other hand, isn't low or mid budget:
Toy Story - $30 Million
A Bug’s Life - $120 Million
Toy Story 2 - $90 Million
Monsters, Inc. - $115 Million
Finding Nemo - $94 Million
The Incredibles - $92 Million
Cars - $120 Million
Ratatouille - $150 Million
WALL-E - $180 Million
Up - $175 Million
Toy Story 3 - $200 Million
Cars 2 - $200 Million
Brave - $185 Million
Monsters University - $200 Million
Inside Out - $175 Million
The Good Dinosaur - $200 Million
Finding Dory - $200 Million
Cars 3 - $175 Million
Coco - $175 Million
Incredibles 2 - $200 Million
Toy Story 4 - $200 Million
Onward - $175 Million
Soul - $150 Million
Luca - Unknown but probably around $150 Million
Turning Red - $175 Million
Lightyear - $200 Million
For that amount of money, they had better pay attention to detail.Miyazaki is doing way more with much less.
Voices of a Distant Star was one person -- Shinkai. That's the kind of thing we'll see more and more of. Small creators reaching audiences and building studios. Gooseworx, psychicpebbles, Vivienne Medrano. That's the algorithm of tomorrow.
AI, as a tool, makes this more possible. One of the first people to do it successfully was Joel Haver, and he's just the first of many to come.
I disagree engineers and artists should do more per unit time. like we need more content per second....
....as if art and real inspiration would ever follow the chaotic beat of human progress