Amazon Polly – Lifelike Text-To-Speech
aws.amazon.com
aws.amazon.com
For a service claiming to be "lifelike", Joey has totally phrased that question as a statement.
[1] https://deepmind.com/blog/wavenet-generative-model-raw-audio...
I don't see (or hear?) the advance here...
Edit: Yes, that example was unfair, because it's from the 90s (and actually, DECtalk was largely unchanged since the 80s, until it was ruined in the late 90s). Here's a somewhat more recent one: With the ETI-Eloquence engine, which was last updated in 2002, the string "2 Marketing" (which one might find, say, in a Wikipedia table of contents) is expanded to "March twond keting".
Someone linked another example here:
https://soundcloud.com/zack-bloom/amazons-new-text-to-speech...
And it pronounces 1970s (the decade) as one-thousand-nine-hundred-seven-ths.
I don't know if this is crossing into uncanny valley territory. I think some of this is not being used to it.
Google Map's driving voice sounds subtly better than that voice sample. I'm looking forward to seeing if a similar service is coming for GCP.
This would drive me nuts, as there's no easy way to predict what sort of transformations it would make. Basically, if it's not near perfect, it's arguably worse than no intelligence at all.
With our PBX stuff, I know it's dumb, so I'm already accustomed to giving it something I know it will read correctly, like "domain dot com".
I believe that's what the Lexicons are for within the service.
I read the post I was replying to as some sort of undefined, built-in, set of intelligent transformations.
For example, 75F-> 75 Fahrenheit would require that you make a Lexicon entry for every value...1F, 2F, 3F, etc. That's not really intelligent :)
Polly 1: https://d0.awsstatic.com/product-marketing/Polly/HelloEnglis...
Polly 2: https://d0.awsstatic.com/product-marketing/Polly/HelloEnglis...
WaveNet 1: https://storage.googleapis.com/deepmind-media/pixie/us-engli...
WaveNet 2: https://storage.googleapis.com/deepmind-media/pixie/us-engli...
There just isn't a comparison -- Polly sounds way more robotic.
WaveNet is Google's, and I'm saying that this is significantly worse.
I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models for tonal languages like Mandarin don't share this problem, since people are speaking with prescribed pitches more often?
Here's a good writeup: http://distill.pub/2016/deconv-checkerboard/
I doubt this would result in audible artifacts at the audio level, though. The visual checkerboarding appears at the 1px detail-level of the output image—this would translate more to minor fluctuations along an audio sample's Nyquist frequency, which is effectively undetectable if the output sampling rate is of the usual kind (44.1kHz).
I mean, a word or phrase could very much change meaning depending on how its spoken (excitement, disgust, fear, sarcasm etc) and even if the context is available in the text (as is often the case), its rather difficult to machine-detect, at least, for now.
Example (click on the blue speaker icon): http://dict.baidu.com/s?wd=%E7%88%B1%E5%B1%8B%E5%8F%8A%E4%B9...
At a secondary level, I imagine there's a sort of pitch trajectory for a sentence, and if a word or syllable ends too far off that trajectory, the return to the pitch trajectory is fast enough to be easily detectable.
It's worth noting that purely synthetic speech, while lacking the superficial lifelike quality, doesn't have these little inconsistencies. Perhaps that's one reason why many blind people, particularly power users, prefer pure synthetic text-to-speech engines, such as ETI-Eloquence (commercial) and eSpeak (open source). These systems especially sound better at high speeds than the newer "lifelike" ones.
Just for fun, here's a side-by-side comparison between a concatenative system, IVONA (acquired by Amazon):
http://mwcampbell.us/tmp/derefr_comment_ivona.mp3
And a purely synthetic system, Eloquence:
http://mwcampbell.us/tmp/derefr_comment_eloquence.mp3
The latter is the text-to-speech engine I use every day, and that clip was synthesized at the speaking rate I normally use.
Edit: It appears the word I was looking for when I said "purely synthetic" is "parametric".
• https://storage.googleapis.com/deepmind-media/pixie/us-engli...
• https://storage.googleapis.com/deepmind-media/pixie/us-engli...
1. How does that engine sound when reading arbitrary text at high speed?
2. Can I run it on my own computer, so I can have low latency in my screen reader? (That, of course, is a problem with Amazon Polly as well.)
This is in comparison with Adobe's tech where it samples a real human voice and allows speech to be inserted (http://arstechnica.com/information-technology/2016/11/adobe-...).
Also, with Amazon's FPGA investment how long before it implements WaveNet using FPGA's to reduce the speech generation time to something that is realtime?
However, Spanish and English are awesome and really sound like a person. This is great!
Edit: After playing a little bit with it, it does sound better than anything I have seen before even in Portuguese. I believe that the weirdness is just for the standard/demo phrase. When I tried some custom phrases (curse words, mainly) it does sound very natural.
One thing to keep in mind regarding speech to text: if you need accuracy, it helps immensely if you can guide the recognizer with a list of expected answers. Un-hinted transcription will do its best to recognize words but without AI to understand the context, there are still quite a few possibilities for mis-recognition. More info about these two here[2].
[2] https://www.tropo.com/docs/voice/transcription-vs-speech-rec...
(Full disclaimer: I'm in the Tropo BU)
This experiment does cover the "generic computer voice" role quite well, but customers don't want that on e.g. their website. I would want a "I'm a professional designer" voice. So there will always be markets for more appropriate voices [*]. But this one does sound very nice at what it does.
The same will happen with robots, btw. Once we get a generic humanoid robot that can do everything, we will still seek aesthetic variations and employ remote body actors.
The idea of applying machine learning to add natural emphasis to speech is a good one, but seeing how language machine learning systems like Google Translate fail on very basic contextual cues, I don't have much hope... Especially when their demo samples are on par with systems from a decade ago.
I'll stick to Bruce.
To play around with it: https://console.aws.amazon.com/polly/home/SynthesizeSpeech
Google, IBM and Microsoft all have invested way more in AI already and are actually way ahead of Amazon technically. However Amazon now becomes the first one that makes these technologies really approachable for ordinary developers. Way to go Amazon!
I remember it sounding pretty life-like. Don't know if they expose an API for it, and sell it though. (It would be great if they did.)
PS.That would be one more job lost to AI.
> The size of the input text can be up to 1500 billed characters (3000 total characters). SSML tags are not counted as billed characters. [1]
Not sure where you got 1000 from. As to the why - no. It doesn't seem like too much. 2 minutes of speech maybe in there?