WaveNet implementation in Keras
github.com
github.com
This shows how far this model is from realtime usage. However I'm sure Deepmind researchers are already looking into how to make this blockbased or some other optimization strategy.
This implementation says: “A Tesla K80 needs around ~4 minutes for generating a second of audio at a sampling rate of 4000hz”, which is significantly faster.
Can someone elaborate on the usefulness of this implementation for Text-to-Speech?
I'm keen to experiment with voice synthesis. I want to create dialog, from multiple voice sources, for some characters in a VR application that I'm working on.
Perhaps this lib is a better option for TTS:
https://github.com/ibab/tensorflow-wavenet
I guess I could do with an ELI5 on how I'd approach this with either of these libraries. I'm not familiar with any deep learning frameworks. But I am pretty handy with Python and have implemented SciKit stuff.
Also thinking this will give me a reason to try Azure K80 instance vs the AWS GPU instances I've been using for other stuff. That said, is a Tesla K80 the only option for WaveNet? I'm guessing I could run it on other GPU's but had read that memory might be an issue on some cards. If so what the lowest card I can run it on and will one of the AWS GPU instances suffice? I also have a GTX 970 at home, but I'm guessing that won't cut it.
I'm looking forward to seeing faster implementations in the future, playing around with this looks like a lot of fun.
Here's a good TTS system:
Further, if you pine for the fjords of DNN-land, merlin (https://github.com/CSTR-Edinburgh/merlin) is brand new and looking to make things a little easier for everybody.
Thanks for the links, but to my ear the samples on those links don't hit the mark. The Wavenet samples in the original article cross the threshold for me. So I'd like to try some short length dialog tests, especially as I've read elsewhere that 1 second only takes 4 minutes on a K80.
Any light anyone else can shed on this would be great.
Looks like I'll have to concede that voice acting is much more practical, for now at least.
In the paper, they say that they double the dilation factor up to a limit and then repeat: 1, 2, 4, ..., 512, 1, 2, 4, ..., 512, 1, 2, 4, ..., 512
The doubling of the dilation factor makes sense to me, but what is happening with the "repeat" part? I don't understand what they are trying go do. Wouldn't make more sense to continue doubling?