Roger Ebert: Hello, this is me speaking
rogerebert.suntimes.com
rogerebert.suntimes.com
It'd be a cool graduate project... kinda wish I was into linguistics right now.
Actually, I think that's what was attempted in the first place in the early 80's. I remember seeing TV shows and museum exhibits that demonstrated this approach. One I especially remember, and highly dates the efforts, had a vector imaging display (think of the original Tempest and Asteroids arcade games) project a silhouette of a tongue and vocal cavity to demonstrate how the current phoneme was generated to listeners.
Of course, back then, such simulacra were limited by lack of parallel processing power and inadequate understanding of biophysics. This lead to the brute force "sound sampling" approach nowadays as memory became more cheap and audio capture hardware was perfected. I do wonder if it's time to return to vocal anatomy modeling again, now that we have a better understanding of how to perform biometric and physics modeling via massive computational parallelism.
The anatomical model does indeed sound very interesting. Each phoneme would be recognized as one particle, on which intonation and dynamics effects could be applied algorithmically; and much of advancement in this area would be employable by speech recognition models, probably increasing their accuracy by a considerable amount.
I really hope some serious contenders step up to the plate for this.
This combined with improved subvocal stuff like http://www.youtube.com/watch?v=xyN4ViZ21N0 would make silent, covert voice communication possible. No more annoying one-sided conversations from cell phones.
It reminded me how far consumer voice synthesis has yet to come, but it did give me a better appreciation of some of the more subtle things the Alex voice has in terms of intonation. Despite still sounding obviously synthetic, it's obviously doing quite a bit of analysis on the sentence structure to vary the pitch in a natural way.
But that makes me wonder: Why, when complex things like structural intonation are already in consumer TTS products, do (deceptively) simple things like consonant sounds and pacing still sound so stilted?