text2speech and speech2speech seem like very different problems no? Especially when so much of someone's voice identity is based on speech patterns, contextual inflection, rhythm/cadence, how would realtime work? Just wondering if you are thinking about realtime as in <10ms or realtime as in 1 second?
Awesome progress by the way, would totally follow along more closely on this project through an eng diary or mailing list.