SAM is awful but at the same time tantalisingly close (one of the demos of it apparently draws on a great reverse engineered version by Sebastian Macke and refactoring efforts by a couple of others including me - I spent too many hours listening to SAM output...) - especially when comparing to the still awful Festival/Flite models -, that I keep wanting to see what a better generic formant synth used as constraint on an ML model would produce.
That is, instead of allowing a generic machine learning model to output unconstrained audio, train it on the basis of letting it produce low bitrate input/control values for a formant synth instead, and see just how small you can push the model.