Most previous attempts at neural net composition restricted the training set to one style of music or even one composer, which is pretty silly if you understand how neural nets work. It was obvious to me if you used a very large network, chose the right input representation, and most importantly used a complete dataset of all available music, you would get great results. That's exactly what OpenAI has done here.
It is still lacking some longer term structure in the music it generates (e.g. ABA form). But I think simply scaling up further (model and dataset size both) could fix that without any breakthroughs. This seems to be OpenAI's bread and butter now: taking existing techniques and scaling them up with a few tweaks. (To be clear, I don't mean to minimize what they've done at all. "Simply" scaling up is not so simple in reality.)
What might still need some breakthroughs is applying the same technique to raw audio instead of MIDI. Perhaps what is needed is a more expressive symbolic representation than MIDI. I'm imagining an architecture with three parts: a transcription network to produce a symbolic representation (perhaps embedding vectors instead of MIDI), something like this MuseNet for the middle, and a synthesis network to translate back to raw audio. This would be analogous to gluing together a speech recognizer, text processing network, and a speech synthesizer. Such a system could generate much more natural sounding music, even perhaps with lyrics.