Algorithmic music composition has usually been split into two:
1. Generate notes (re: theory, genre)
2. Generate sound
(i.e., EMI[0], Kulitta[1], MusicNet[2])
Now we are doing both at the same time, and backwards. The model isn't (necessarily) going "write melody, then generate the sound", but rather, "here are 500 songs that are described with X, 500 with Y, and you want XY, so we'll combine these two" :)
(This is my best understanding, so feel free to correct)
[0]: http://artsites.ucsc.edu/faculty/cope/experiments.htm