An open source implementation of DeepVoice 3: 2000-Speaker Neural Text-to-Speech
r9y9.github.io
r9y9.github.io
Post-processing by convolving with a short noise burst removes them pretty well, as that randomizes phase vs. frequency.
https://drive.google.com/file/d/1WwmCNwWwukhtYXRT4mDHxCUa0sQ...
But it also makes it sound like she's in a closet, so it would be better to fix it at source if there's a way to do that.
https://google.github.io/tacotron/publications/global_style_...
Those samples sound already frighteningly human.
Not to disparage the DeepVoice3 in anyway! There must have been a lot of work put into it.
Funny how they publish their "research papers", yet no one else is able to implement their engine with even remotely comparable results.
https://www.nuance.com/en-gb/omni-channel-customer-engagemen...
Just know that the voice will be similar to what Kyubyong or others managed to train meaning it will sound eerily synthetic. Might fit your purposes but it's probably not enough for consumer-applications. Also from what I played around with it optimizing the synthesization is going to be big hurdle if you want it done quickly. Or not I didn't dig that deep into it.
https://cloudplatform.googleblog.com/2018/03/introducing-Clo...
Suppose to be using 16k samples a second through a NN which seems hard to believe. But gets you a pretty incredible result.
It seems to be all Pytorch in the code though.
https://machinelearning.apple.com/2017/08/06/siri-voices.htm...