Algorithmic music composition has usually been split into two:
1. Generate notes (re: theory, genre)
2. Generate sound
(i.e., EMI[0], Kulitta[1], MusicNet[2])
Now we are doing both at the same time, and backwards. The model isn't (necessarily) going "write melody, then generate the sound", but rather, "here are 500 songs that are described with X, 500 with Y, and you want XY, so we'll combine these two" :)
(This is my best understanding, so feel free to correct)
[0]: http://artsites.ucsc.edu/faculty/cope/experiments.htm
The problem is that coherent musical structures are much more constrained. You can't just XYZ... into a space and get something that makes sense.
That will kind of work for low-density music, which includes a lot of landfill dance + subgenres. But these statistical models are blind to larger and more complex structures, and completely unaware of cultural context and semantics.
It's actually a harder problem than language modelling because the spaces and the grammars are much larger, especially once you start including sound quality and production values as well as arrangement and core composition.
We are already drowning in music, you can turn on Spotify and have enough music to fill a lifetime. Yet new music is still being produced, why? Because music is ultimately a psychological experience, the human connection is a not-insubstantial part of the experience.
There’s a place for AI in music but it has to be white box, there needs to be scope for a human to jump in there, modify things, and make it their own. Otherwise, who will care?
And an infinite stream of music is exactly what I want. I don't want to curate or search. I want to feel and ask and get.
FWIW, the various diffusion models (tend to) use cross attention to an attention-based approach.