Maybe Tacotron will interest you?
It's an end-to-end model, that's reasonably close to the state of the art:
https://google.github.io/tacotron/publications/tacotron/inde...
They are some open source implementations.
https://google.github.io/tacotron/publications/tacotron/inde...
They are some open source implementations.
Edit: Another interesting one: http://research.baidu.com/deep-voice-3-2000-speaker-neural-t...