Google’s DeepMind Achieves Speech-Generation Breakthrough
bloomberg.com
bloomberg.com
And the WaveNet site with audio samples: https://deepmind.com/blog/wavenet-generative-model-raw-audio...
The comparison against state of the art Parametric and Concatenative methods are pretty mind blowing.
Particularly listen to the music samples. That's a generated piano piece that sounds quite musical.
They even include breaths and other auditory signals that really make for a convincing speech sample.
That blew my mind. I've heard computer generated music before, but it's been just synthesized with usual methods while the computer is just the "composer".
I find the crackling and buzzing a bit awkward, though. It's probably an artifact of the algorithm and can probably be mitigated with some simple filtering.
Fixing that requires higher than sentence-level learning, since it has to know that it's introducing a potentially unfamiliar term which can only be known from the sentence's context within the article.
Do you remember a film titled Rising Sun?
Or, alternatively, we do not go that route and as a result the burden of truth is diminished to the point where factual information will no longer be enough to make a definitive judgement.
This Pandora's Box is staying open. Let's all hope for the sake of civilization that we use its powers responsibly.
i figured knowing obama's interest in good tv shows, he might have done the cameo! this is really neat stuff
Reminds me of "virtual" kidnapping scams where an attacker knows the victims phone will be unreachable and the attacker calls a relative demanding they wire funds or the victim will be killed. Attacker plays back a voice sample they've captured from the victim that makes it sound like they're in trouble an need help; basically the attacker calls the victim and say something like "May I help you?" repeatively until the victim responds with something like "No, I don't need your help!" an panicked voice - which is then edited to say "Help! I need your help! Help!"
https://deepmind.com/blog/wavenet-generative-model-raw-audio...
David Attenborough, Morgan Freeman and Billy West: call your offices.
Recording: https://youtu.be/2ljFfL-mL70
Quoting from the episode: "Essentially what it says is that you can't take somebody's name, or their likeness, or their signature or their voice, and use it for commercial gain without permission from that person. And if you do that, you're liable."
The full episode is here - it's one of Planet Money's best. http://www.npr.org/templates/transcript/transcript.php?story...
- Morgan Freeman
- Alan Rickman
- Liam Neeson
- David Attenborough
- John O'Hurley
- Denzel Washington
https://deepmind.com/blog/wavenet-generative-model-raw-audio...
16,000 analyses per second. So, given Moore's law, around the iphone 9?
That's a lot of cranks of Moore's law. Better hope for considerable algorithmic improvements. (Raw is probably overkill anyway.)
Not stating anything especially mind-blowing. Just restating the shocking speed of exponential growth/shrinkage.