out of interest, do you also work on a reverse solution, text-to-speech? Most open source engines sadly still can't compete with commercial alternatives.
https://google.github.io/tacotron/publications/tacotron/inde...
They are some open source implementations.
Edit: Another interesting one: http://research.baidu.com/deep-voice-3-2000-speaker-neural-t...