Recurrent Neural Nets for Speech Synthesis
arxiv.org
arxiv.org
http://www.zhizheng.org/demo/is15_mte/demo.html http://www.zhizheng.org/demo/dnn_tts/demo.html
SPSS is the problem of going from linguistic features, phenomes, etc, to speech audio. These features are more or less golden, either derived from the audio itself or hand-labeled. So things like tonality, cadence, emphasis on words is already encoded as features which is why these samples sound so good.
Deriving these features from pure text is very hard, and this failing is the main reason most text to speech systems sound so dull and tone-dead.
That being said, these results are seriously impressive, sounding very natural. Would love to see someone try and train an end-to-end system from pure text to speech. I think we'd see some big improvements like what Baidu has done for end-to-end speech to text.
There were a few efforts to make actual silicon neurons, plus the whole nueromorphic movement, but they were generally less than what people were expecting, slow, and difficult to interface with.
That work seems to spin their contribution as reducing the power required to evaluate the neural network though. If I recall correctly, the accuracy of those models for everyday tasks is typically much much lower than usual ANNs, and they're a pain to train. So, still not very common.
1:"Hey, do you see the first squiggle with the two fuzzes after."
2:"Next to Beaker's eyebrows?"
The low power work seems to have been aiming to be a rough filter, rather than a full system. Still fun to use.