Char2Wav: End-To-End Speech Synthesis
mila.umontreal.ca
mila.umontreal.ca
It leads to "AI hallucination", where even if "foo2bar" doesn't work, people assume that it's the one right AI for turning foo into bar. When someone gets better at turning foo into bar, the typical response will be "is that just foo2bar?"
This happened absurdly backwards with doc2vec, which after word2vec everyone talked about as if it were a real thing, until Radim Řehůřek finally made a reasonable implementation of it under that name.
NN is universal function approximator, X2Y is just what it is.
Describing everything as a "universal function approximator" is misleading if you never look into how good the approximation is. Char2wav, for example, is a neat trick, but clearly wouldn't be used for real speech synthesis.
See some early work on Romanian (https://www.youtube.com/watch?v=cwnDjq33uMs), and compare to Google Translate TTS (the second to last, robotic example). The existing TTS systems out there in the research community are also pretty good (last example), but I think we are at least competitive which is interesting, given that we are far from TTS experts especially in all the NLP processing that normally happens in building these systems (https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/qu... for example).
At very least the Edinburgh system (http://romaniantts.com/new/) should be the baseline across companies for Romanian - but one of the things we also want to show is that one architecture/approach can generalize easily across several languages. By and large the approach for all our languages is identical, including nearly all hyperparameters (we change one setting for attention default step size, and I think that is it).
There are tons of languages which have poor existing systems, and having something that is basically - record a speaker for a while, write down whatever sentences they are saying, train a big model, and do pretty well - could be very useful. We are still exploring what languages this approach works for, but have had a good success rate so far if we can find a good openly available dataset. It also opens the door to using existing sentence level ASR datasets "in reverse", since we don't need timing/alignment information.
There is also a lot of potential for personalized sound/speakers using speaker interpolation (see the alternate speaker examples in the youtube video) that we have not explored yet, as well as applications to related sequence generation tasks. I think some of our training tricks are useful for training sequence generation models in general.
Some great videos that help put our work in context [1][2].
[0] early Romanian Demo of our approach: https://www.youtube.com/watch?v=cwnDjq33uMs
[1] Alex Graves, Generating Sequences with Recurrent Neural Networks (see ~36m in for speech demo): https://www.youtube.com/watch?v=-yX1SYeDHbg
[2] Heiga Zen, Generative Model-Based TTS Synthesis: https://www.youtube.com/watch?v=nsrSrYtKkT8
I have not laughed this hard in a long time.
for example, on this page, only spanish has the char2wav label.
http://www.josesotelo.com/speechsynthesis/
It's unclear which results are the output of the model.
Also, I notice that many of the result clips trail off in volume. Is that a processing error or intentional in how the clips are edited?
Regarding the truncation at the end, that was a bug in our sampling code that we just fixed. We will update the samples soon!