Amazon Polly – Text to Speech in 47 Voices and 24 Languages
aws.amazon.com
aws.amazon.com
You can read this as: We're pleased to announce our really cool TTS feature that took a lot of engineering know-how and effort ... but you'll have to click through, because we can't seem to get around the limitations of our CMS to embed audio content in a blog post.
I realized this tool might not be cheap, since it may take the voice actor/actress 2 hours per day to produce my content (2-hour driving commuting per day for me). To get familiar local accent, it costs ~$36 in Australia, and maybe slightly cheaper for US accent. The value it brings me can hardly justify the cost.
Now, with Polly, things changed - it produces reasonable voice, and 2-hour content would only cost ~$0.3. I decided to launch my service as soon as Instapaper approves my API request.
At the same time, put your email here: http://readlater.launchrock.co/
EDIT: It is robot-like and nowhere near the quality of the samples that Amazon provided, but it gets the job done.
My biggest problem with CereVoice, though, has been its terrible web API. It doesn't support streaming output, so it renders the audio to an Amazon S3 bucket and then returns a URL, which is pretty inconvenient (and slow). You have to do the same for transcripts, too. So, if you want everything, you have to make 3 separate HTTP requests and parse 2 XML documents for one round of synthesis.
IBM Watson's TTS API gets it right, imo. Its streaming mode returns audio frames and transcripts over a WebSocket connection.
https://deepmind.com/blog/wavenet-generative-model-raw-audio...
Computer generated voices feel most robotic when their intonation of a word is abnormal or their pauses between words make the sentence feel choppy. The intonation and natural pauses between words is very good for all of the main voices.
The Japanese voice Mizuki was the most comical addition, since I can't think of a real situation where she would ever actually be used. Mizuki speaks Engarish (the Japanese version of English) beautifully, but any Japanese person who can understand Engarish will also understand English. Also, Mizuki doesn't add the correct vowel ending to all words, e.g. she correctly says "cheezu" for "cheese", but says "steku" instead of "steki" for "steak".
My impression is that it's actually quite common for Japanese people to understand Japanese accented English more easily than a native English speaker.
>she correctly says "cheezu" for "cheese", but says "steku" instead of "steki" for "steak".
"Suteeku" sounds like what a Japanese person who knows English well but has a strong accent might say. "Suteeki" is more of a corrupted loanword.
This is actually very true, made me think back to many situations in Japan where adding a strong Japanese accent to my English words made it comprehensible to the listener (just as adding a Japanese accent to the generated speech makes it nearly incomprehensible to a non-Japanese speaker).
It'd be tolerable to hear a voice-interface with this, but it'd be maddening to try to listen to a book this way.
> F. Voices. Subject to the terms and conditions of this License, you may use the system voices included in the Apple Software (“System Voices”) (i) while running the Apple Software and (ii) to create your own original content and projects for your personal, non-commercial use. No other use of the System Voices is permitted by this License, including but not limited to the use, reproduction, display, performance, recording, publishing or redistribution of any of the System Voices in a profit, non-profit, public sharing or commercial context.
Which truly sucks, because they don't even give you the option to pay for such use.
https://d0.awsstatic.com/product-marketing/Polly/pt_br_vitor...
Unfortunately, I agree with Mizza regarding the quality.
EDIT: oh, I had no idea they have used Ivona
edit: ah it was just the german version that is not available any more, english one seems to be still the store: https://play.google.com/store/apps/details?id=com.ivona.tts
It really is quite good--even if I really wouldn't want to read an entire book read this way or would mistake it for a human. It definitely gets me thinking about ways to use this service.
English male sounds surprisingly bad. English female is better.
http://g-ecx.images-amazon.com/images/G/01/ivona/static//med...
I hoped there would be new & better german voices
This could get a bit confusing for some folks.