They really need to bump up to 48 KHz so all the music doesn't sound like it's being played over a telephone. A factor of two in the cost shouldn't be prohibitive. So much of the audio generation stuff I've seen has fatal flaws like this baked into the dataset and/or training process that ensure the output can't sound good even in theory, it's kinda frustrating.
It's also frustrating that nobody AFAIK has trained a music model on an actually large dataset. We're training large language models on a significant fraction of the text of the whole internet! Where are the audio models trained on a significant fraction of all recorded music? This one was trained (in part) on a dataset of 280k hours of music, if I read the paper correctly. I don't think that comes anywhere close to being a significant fraction of all recorded music.
https://ooo.ghostbows.ooo/about/
Music is on all the usual streaming platforms, the album is called Shadow Planet, band is The Cotton Modules. Also avail on their website:
There is a million Industrial techno sounds, repetitive, hypnotic rhythms IDM tracks that absolutely no one listens to anymore. Most the founders of the whole genre have moved on from lack of interest.
We have been able to do really good generative music in Reaktor for the last 20 years. Much better than anything on this page but no one cares.
People in general want to hear the same slight variation in music over and over as background noise.
The people that are saying how great all this is will be bored with it in 2 weeks or less. Pure technological kitsch.
These models are still in their early phases - If AI was combustion engines, we wouldn't even be to railroads yet. Despite that, they're still fun to play with and useful when used correctly.
I was born too soon to explore the stars, born too late to be a 17th century pirate... but born just in time to explore these incredible neural networks come to life.