As someone who's worked a lot on procedural music, I think this is definitely true. I'm always surprised to see ML-based approaches where someone has just trained a system on a bunch of songs, and then hopes the system will produce music with a recognizable structure - even though all the training songs will have had (in general) different chord progressions, different numbers of voices or melodic lines at any given time, etc. Such approaches strike me as akin to training a system on a bunch of short stories and then hoping it will produce a new story with a recognizable plot.
It seems like it would make a lot more sense to remove these hidden dimensionalities, e.g. by annotating the source data with chord or structural information, or by training on lots of different melodies that all share the same chord progression, etc. But it's hard to imagine that with enough layers the network will eventually grok all these hidden details.