Making a TTS model with 1 minute of speech samples within 10 minutes
github.com
github.com
If trained on longer samples does it start to hallucinate? Learning programs are notorious for rough approximations failing to scale into useful detail.
Very cool work though!
The paper didn't mention normalization, but without normalization I couldn't get it to work. So I added layer normalization.
The paper fixed the learning rate to 0.001, but it didn't work for me. So I decayed it.
I tried to train Text2Mel and SSRN simultaneously, but it didn't work. I guess separating those two networks mitigates the burden of training.
The authors claimed that the model can be trained within a day, but unfortunately the luck was not mine. However obviously this is much faster than Tacotron as it uses only convolution layers.
The paper didn't mention dropouts. I applied them as I believe it helps for regularization.
I heard a Google TTS demo (I think it was called deepmind?) that sounds extremely human-like and I was wondering if it can be used to turn webpages into speech (I have few extension in Chrome but those voices sound very robotic and it's hard to hear them after 2 mins).
Anyway congrats for making this. I'm nowhere near as smart to understand how it works just that it's getting better and more human like everyday!
The technology is definitely intended for reading websites to users. I'm actually not sure why Google hasn't integrated it into Chrome yet. Maybe they prefer leaving the task to the OS-level accessibility tools.
They couldn't afford for everyone to be text-to-speaking every web page...