Short answer:
Don't use this for practical purposes. It takes 90 minutes to generate 1 second of audio.
Here's a good TTS system:
Here's a good TTS system:
Further, if you pine for the fjords of DNN-land, merlin (https://github.com/CSTR-Edinburgh/merlin) is brand new and looking to make things a little easier for everybody.
Thanks for the links, but to my ear the samples on those links don't hit the mark. The Wavenet samples in the original article cross the threshold for me. So I'd like to try some short length dialog tests, especially as I've read elsewhere that 1 second only takes 4 minutes on a K80.
Any light anyone else can shed on this would be great.
Looks like I'll have to concede that voice acting is much more practical, for now at least.