If you can use a piece as inspiration, browsing around until you find the 5% or whatever % that you like, there's really some benefit to be found. I felt a bit like Simon Cowell, auditioning works and being really picky, last time I did this. But in the end you can discard an item, go with it as-is, adapt it somehow, or use it as reference material. Eventually you build a gallery.
Do people ever train at different hierarchical levels? I've done very little ml, but it seems to me that it'd be beneficial to train a net on "plans" and then separately train one to interpret plans.
Music has many rules, not only theoretical but unwritten rules about what works. You must incorporate them somehow into the program, either by code or offering something to the program to deduce them.
At the moment, it probably would be more practical to train a model to predict ratings and use that to screen generated samples or possible completions and throw out too-low-scoring ones (the 'ranker' approach worked out very well for the Meena chatbot recently: https://arxiv.org/abs/2001.09977 )
Music falls somewhere in between text (as a sequence of chords or PCM samples) and image (as a piano roll or a spectrogram), so maybe some hybrid of image and text generators is needed.
If you wanted to improve my ABC-MIDI GPT-2, the most straightforward ways would be to do data cleaning (I'm sure there's tens of thousands of awful MIDI files which should be removed! data cleaning with RNNs or GPT-2 or GANs always makes a large difference) and increase the model size (the fact that loss bottomed at 0.20, which is still quite bad, suggests that MIDI is hard enough that GPT-2 is struggling). More interesting would be to use Reformer or another long-range Transformer and try to operate directly on a more raw representation, like the the piano roll representation of MIDI. I think GPT-2 makes a lot of syntax errors which cripple outputs when a 'voice' goes silent, and a piano roll representation would be a lot more robust (at the cost of being like 10x larger).
Personally, I'd rather try redoing BPE encoding for ABC-MIDI specifically.
It doesn't, incidentally. It's just a standard StyleGAN dumping random images, trained to to model the average image/distribution, and not optimizing for human ratings or anything. Almost all the 'X Does Not Exist' things operate that way, including my own https://www.thiswaifudoesnotexist.net/