can we eventually apply style transfer to speech? I'd like to hear all audio narration in the voice of David Attenborough.
Edit: It would seem they only need 25 hours of audio to train the model, so it would probably be easy to find or make alternate datasets.
> We train Tacotron on an internal North American English dataset, which contains about 24.6 hours of speech data spoken by a professional female speaker
In the case of David Attenborough, I'm pretty sure that amount of data is available with subtitled BBC content.
I honestly wouldn't be surprised if you could find a really feature rich sentence (say 5 words long) that you could use to crack pretty much all voice-activated biometric password systems.
[1] http://www.washingtonpost.com/wp-srv/national/dotmil/arkin02...