I know nothing in text to speech, so maybe my question is stupid, but I've always wondered if somebody tried to produce "natural" sound by modelizing the air through a human mouth+throat+nose, so that your would have the "naturalness" of the voice (especially if you add the dynamics part, like air volume in the lungs that force you to pause, time it takes to reposition the tongue/mouth between two sounds) , or if it's actually more complicated than that/too ressource heavy/too hard to modelize etc.