Awesome work! One thing I'm curious about in this space is why people generally generate the sound form directly. I always imagined you'd get better results teaching the model to output parameters which you could feed into synths (wavetable/fm/granular/VA), samplers, and effects, alongside MIDI.
You'd imagine you could estimate most music with this with less compute and higher determinism and introspection. Is it because there isn't enough training data for the above?