WaveNet launches in the Google Assistant
deepmind.com
deepmind.com
I'm sure things have improved significantly since 2012 so what you're looking for is probably easily done.
It can read .epub files directly, along with text files and webpages, and integrates with Dropbox, Pocket, Gutenberg, etc.
On Android, you can get Voice Dream Reader (though it has less features than on iOS) or @Voice Aloud Reader which is free [2].
However, it was unusable because it would stop as soon as the screen would turn-off. Other ebook readers' read-aloud features worked with the screen off. So thanks for nothing Google!
Some of the voices that came out of that were incredible.
Also music generation - the original post (September of last year) had some really interesting music created by training the system on Chopin or something.
I am the tech manager for a machine learning team and potentially eliminating jobs is the bad side of my field. That said, I think AI, in some sense of the term, will also help us be more efficient. I imagine everyone having effective assistants that help us in our work, help with communications and scheduling, etc.
Uses Amazon Polly for the speech gen
http://www.voicedream.com/reader/
It has so many voices available in the voice store, and I've had no trouble reading books with it. I'm constantly amazed by the quality of some of these voices when played at high speed... They take breaths, express emotion, everything. My only problem is they occasionally mispronounce brands, names & acronyms: Quite amusing to learn about "Sharr-E-poynt" instead of "SharePoint", "Amd" not "A.M.D", or use an incorrect pronunciation of a character in a book when talking about the book... Can't really see how WaveNet would fix that, but look forward to seeing these voices turn up in Voice Dream.
There are noticeable blips in the speech that sound unnatural, particularly when certain sound combinations are used.
The very first sample with "Bruce Frederick" is clearly off. The intonation and timing between the end of bruce and the beginning of frederick is... mechanical.
There's a similar problem in the OPs link with the non-wavenet English voice 1 when it says "Wavenet".
Those issues are much less apparent in the wavenet voices. Timing problems are less noticeable, intonation problems are less noticeable.
Frankly, the voices there sound VERY good, compared to anything I've heard.
That said, I completely agree that there's not enough samples there to make any real judgement.
Apple's change in voice talent is an improvement though, and they may have more units than before, which is helpful. I believe their model also works offline, which is a huge plus (though I think Google's prior model works offline as well).
I learned a useful new word, thank you!
Still, they picked one that makes theirs look vastly superior.
Listen to the samples here for example: http://voicetext.jp/samplevoice/
I like the sound of Risa in http://voicetext.jp better. Try it out the Japanese Google used in their example. If someone knows of a better one it would be great to find out.
一つのウエーブネットのみで、多数の事なら、話者の音声波形極めて正確モデルができます。
Also, for traditional systems you need a lot of data from one speaker only, they can't take advantage of other speakers' recordings (although WaveNet does that now).
And "TV natural" might not be the style of natural you want from a TTS system.
Google, to their credit, normally allows you to erase all this data and turn recording off, but with Google Assistant, they require it all to be on and recording. I avoid using it because of this restriction.
I did try out Google Now for a time due to having an Android Wear watch and found that as time went on more functionality required the turning on of these settings. One particular oddity for a time was that "OK Google, Navigate Home" would produce only a complaint about web and app activity being off, but "OK Google, navigate to $HOME_STREET, $HOME_TOWN" was fine.
I've abandoned Android Wear because of this.
I wonder what compromises they made to improve the performance by 1,000x. There have to be some.
WaveNet is probably modeling the source data very well. It sounds like they just need more data with emotion and inflection, rather than having source data that is optimized for monotonicity and precision.
If you work in radio or voiceover you learn really quickly that the voice is so much more complex than people give it credit for. Subtle changes in delivery, timing, inflection, syllabic emphasis, pauses etc... make a massive difference.
Anyone can "talk like a robot" but speaking naturally is way more dynamical than just making the sounds of the words transition smoothly.
I'm not sure if it's way easier or way harder than we're doing it now.
EDIT: A big example of that is the Japanese example at the bottom. English has had far more effort put into the old model, so it's already pretty good. But the difference between the old and new Japanese voice is really striking, and they were most likely able to make the new one much faster.
You could vary so much
Would e.g. Don LaFontaine's family go all J.R.R. Tolkien and forbid the use of his voice in certain contexts?
Then again, for now, it seems that it's easy to get away with synthesizing voices of popular characters as long as you don't use copyrighted names:
https://acapela-box.com/AcaBox/index.php
(choose English(USA) - "Little Creature". Borked on Firefox but works in Chrome)
"Little Creature" my ass.
Apparently, names are IP, voices not so much. But we can't have nice things because like you said - eventually, someone's going to figure out how to file an effective lawsuit.
For example, should a lifelike digital model (or even a still image) of an actor's face still generate royalties for the actor's family after their death? Both cases are using a unique attribute of the person to promote something.
Google is only having any success with this because they are getting training data from the public for free, they should also return the model to the public for free (at least for noncommercial use)
(And I don't think WaveNet is even based on public training data.)
Public Money, Public Code should also extend to Public Data, Public Code.
The amount of tweaking and finessing that goes on under the hood for specific scenarios can change outcomes quite a lot.
And from a product perspective, it can be taken quite far - for example, most of the 'common' things Siri says are not synthesized, it's literally are recording of the voice over artists. The more arcane stuff is synthesized.
It's always comparing apples to oranges to bananas unless you really know what they're doing, even then it's hard.
Hell, I'd settle for an API where I could send text and get high-quality voice back. Maybe I can somehow hack the Google Assistant app to do it...
In fact that's the title of the article.
Edit: I might be wrong, at the end of a paragraph they say it runs on Google’s TPU cloud infrastructure, though it isn't clear to me whether they just use that for training.
Edit 2: I just tried it on my phone. At least stuff like asking it to "Turn on WiFi" works without an internet connection, and yields a TTS response.
But this is the status quo. You would not expect Google to disable offline TTS just for slightly improved quality. The real question is, is it running Wavenet offline or the previous version of its TTS engine offline?
I haven't actually done much about it yet, but I'm interested.
We had launched a Bot which gives voice summary of web content on Messenger, Slack, Telegram & Twitter with In-line audio player on first three. It’s great for sharing audio summary to our visually impaired friends.
Check it out here - https://larynx.io/#larynxBot
I'd like to be able to record my voice - and have it translated to text, then compare how it sounds via both engines vs my voice/cadence.