It seems like it would be, right? There are 44 tract-diameters you can modify to shape the vocal tract, and these can be used to generated specific vowel formants. I can imagine you can build a system using deep learning that can find the best parameters to match a steady state periodic pitch. It's a bit how some speech codecs work, like LPC10.