Deepfake Voice Technology: The Good. The Bad. The Future
econotimes.com
econotimes.com
Audio deepfakes, when they finally arrive for real, will simply reduce the effort required for voice impersonation. They do not make possible new feats of impersonation that were never possible before. They can still have an effect on society for sure, but I believe the concern is out of proportion with the effect.
This is a global issue. For examples, please look into the societal concerns surrounding 'deep fakes', and 'fake news' for answers to your first two questions.
I don't believe the concerns about deep fakes and fake news. I think they're no worse than what we've always had. Those fears are exaggerated by arrogant hand-wringers who somehow assume their own beliefs are exempt from manipulation and it's only intellectually inferior "other" people who are susceptible. They're so fixated on the popular political ideas that they forget they should be on a crusade against religion which is the elephant in the room when it comes to people being misled.
https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice...
Probably it' easier because phone calls aren't high quality.
I wouldn’t call the voices unconvincing. They used to have a better demo though where it would impersonate celebrities like Trump. It was very interesting to demo it.
there is a difference in scale and effect between every idiot in the world can fake as many impersonations of people that they want (and they can coordinate this impersonation with 100s of others who have a goal with this impersonation) and Rich Little calling up somebody's granny and claiming to be Jimmy Stewart.
It's pretty trash, but it's not completely horrible. It's also over a year old, so the SOTA is likely better today.
I believe the contrary: that high-quality, convincing audio (and video) deepfakes will be cataclysmic for society. Voice audio is one of the last digital artifacts that people will believe with their senses.
When Donald Trump’s “grab them by the p—“ audioclip came out, he was at least forced to acknowledge it.
Once high-quality audio deepfake tools become available, amateurs will be able to flood the web with fake audioclips of politicians saying things they never did, and truly unscrupulous politicians will be given technological cover for any damning audioclips of theirs that get leaked.
Once we cannot trust our senses, then we will be only able to trust our institutions and then our tribes. But trust in institutions is already decaying rapidly, so in the end there only ideological tribes will be left.
Synthetic is scalable.
So you could still deep fake voice actors with the intention of just saving money rather than wanting people to recognise them.
I'd go much farther than that. Unique features of someone's voice and appearance (face, walk) should not only be protected by copyright but also by privacy laws. It would be crazy if producers were allowed to make arbitrary videos with anyone's appearance and voice, putting words in your mouth you'd never endorse.
Luckily, this may be one of the cases when politics will react fast and swiftly, if necessary with new laws, since politicians would likely be among the first victims of the new trend.
....what if you had a malicious video conferencing service that was free to all?
Because it's free it gains mass adoption,the service can target a person it is interested in and can then record all audio and video footage.
Then using this captured data train a real-time puppet.
You could then intercept any call that person is on and 'take over' and push your own agenda, or call people up and do the same, pretending to be that person.
With something like that it would add more value in ongoing OTP ways of verifying a person is who they say they are.
https://www.youtube.com/watch?v=t5yw5cR79VA&feature=emb_logo
Deep fakes audio sounds better than most text to speech models that are freely available.