One of the things I've long wanted to do but not found time for, is to take a few different variants of formant synths, and try to train a simple TTS model to control one instead of producing "raw" output. It's amazing what TTS models can do with "raw" output, but we know our brains aren't producing raw, unconstrained digital audio, and so I think there's a lot of potential to understanding more and simplifying if you train models constrained to produce outputs we know ought to be sufficient, and push their size as far down as we can.