“I Used AI to Clone My Voice and Trick My Mom into Thinking It Was Me”
buzzfeednews.com
buzzfeednews.com
Hoping that she’d one day beat the cancer, but may not have a voice, I came up with an idea of trying to “capture it” in 2009 - hoping that it could be algorithmically rebuilt in the future. I reached out to a number of individuals that ultimately put me in touch with a research group that had a proprietary setup for capturing samples and rebuilding the voice. Over the Thanksgiving break, I managed to get access to a soundproof recording room and they worked with my wife to capture samples over a period of 4 hours.
Having worked in the infosec space since the 90s, my first reaction is often either how new tech/innovation can be used to bypass a control and how one could detect/prevent that. It’s easy to lose sight of how something like this could fundamentally changes a persons life.
Thinking more about the specific use-case you have in mind, I find myself wondering how sentiment and inflection might be captured via a synthetic voice. Would it be inferred by context? How would that inference deal with things like sarcasm/irony. I wonder if there could be some input mechanism for controlling the inflection - what would that input interface look like? Could it go off facial expression?
I wonder where the existing tech sits in the uncanny valley for this space...
I was very confused how she would pass a bowl of unidentified ash from the airport security (we only had a backpack each). She drafted a poorly done and obviously fake death certificate. It was not campfire ash anymore, it was the remains of her father.
The people at the airport were visibly awkward, they tried to be as accommodating as they could. She flew back home with a plastic bowl of ashes from our campfire, it even had some parts of birch and branches.
Airport security was easily fooled. And the author's mom is easily fooled too, motherly instincts be damned. Would a neural net be fooled by the author's attempts? I know for sure that an automated security system would sound the alarm on my friend. I 'd like to see adversarial networks fighting each other on such premises. A son network trying to fool the mother network and vice versa ad infinitum, at least 1 billion of simulation hours in. What kind of wonders would come out
In reality the engines are constantly producing a stream of hot compressed air which is bleed off for various subsystems. One of these is the air conditioning systems. These cool that air and filter it so that it's clean and at a reasonable temperature for passenger comfort. The air is added to the cabin at a pretty much fixed rate and the pressure is regulated by a dump valve which dumps excess pressure overboard. There is no real recycling of air.
How do they regulate oxygen levels?
Or are those levels more or less the same as on the ground but it's the pressure that incommodes?
Questionable items with no obvious cash value simply go away. Off the plane.
It probably is the smell, the charcoal and trace chemicals that mess with swabs and dogs.
They were ... not very happy when we realized what had happened.
I do recommend watching the episode of Follow This as suggested in the article (episode 7) if you’re interested in the latest deepfake tech, and its implications for fooling people who can’t obviously tell it’s fake.
I did not work on Follow This, although I do appear in a B-roll shot. https://twitter.com/minimaxir/status/1034109759295647745
> required
Is this even legal according to GDPR?
https://screenshots.firefox.com/MnEgMtsGavMlxcts/www.buzzfee...
“Sure honey.”
“200 OK. Eh I mean, thanks.”
I have a stutter that is especially bad when I first talk on the phone. I used to do something similar where I would record an introduction and then play it when the phone connected.
The quality wasn't great, but it was better than me not being able to say anything.
So it's impressive, but let's not get ahead of ourselves here.
[1] https://www.youtube.com/channel/UCEOXxzW2vU0P-0THehuIIeg
I want to use it, but I wouldn't use it for any real products today. Amazon Poly isn't amazing either, but sounds more natural than these samples. Yes, it's only a few stock voices, but it's a lot closer.
You know, voice call sound quality go full terrible in real life too. If, say, one of my ex-gf or even my sister suddenly call and sounded a bit robotic like this I would believe them to be who they claim to be, no contest.
Do you believe this isn't an advertisement for Lyrebird?
Like I said, it's a really cool idea. If I could realistically duplicate my voice I'd certainly use that. But as of right now... meh.
As realtime, realistic voice synthesize is thought to be difficult, a voice/phone call encryption system usually circumvents this problem by using the caller's voice as the proof, as both recognize each other's voice. In most phone encryption systems, like many commercial systems, or the ZRTP protocol by Phil Zimmermann, or the "safety number" in Signal, they allow both parties to read out their pubkey's SHA-256 hash digest (usually encoded to words) aloud, as a mean of verification.
If this type of AI-based voice synthesizer becomes widely-available, it could be disastrous to cryptography. It is not the end-of-the-world of course, as targeted attack with social engineering is not an issue for most people, and those who need this level of security is going to perform out-of-band exchange anyway, but still, the certainty of voice-based key verification would be greatly weakened.
Cryptography doesn't solve political problems.
Cryptography solves communication problems. That's. It.
Or am I missing something?
(I say nothing for or against the particular scheme proposed by gp, just against parent's implied generalized dismissal of crypto to solve this problem)
While I understand the parent comment's stance that cryptography is not the magic sauce, I don't think it's related to my comment.
It didn't sound like the mom was buying it either. Her tone of voice was somewhere between "I'm going to play along with this" to "Dear god I think my son is on drugs again."
Not judging his work here, but it sounds like an unstable VoIP call.
In this case, your brain (the "non-artificial intelligence") can take some text and control your vocal chords to emit sound waves to produce speech. You can even learn different voices like a cartoon character voice artist. The artificial intelligence can learn to do the same thing.
Simple AI accelerated and deskilled the former, and combined with ubiquitous social networks exponentially empowered the latter.
The thing is: these truth undermining technologies need not be perfect to have the effect, just 'good enough' to cast significant doubt and allow near everyone to believe their own 'truths'.
The result will be a highly dis-empathic society, where trust beyond the most closest 'clan' is close to nil and even then some.
Confusion always empowered narcissist and sociopaths, the con-artists and the cultists. It isn't so hard to see anymore how old civilizations could devolve into the dark ages.
We have already seen this work on "lo-fi" ads.
Ads with obvious spelling and grammar errors, meant to immediately engage those who have no concerns of such things (filter out the critical thinkers).
http://jamiedupree.blog.wsbradio.com/2018/06/18/back-on-the-...
When you talk to your Mum on any modern phone system, there's an argument that you're actually speaking to a voice synthesiser that sounds like her. Maybe we need to view the thought process as the person rather than the voice?
I always wondered why emails were evidence. They're just text, anyone could fake an email.
The right question isn't usually if they could, but how likely it is they would. Also, I do hope cases aren't usually hinged on a single e-mail, but include other evidence that together paints a coherent story.
It turns out that the metadata footprint left on a computer creating a document- if you can seize it- is rich in detail enabling creation date and the like to be identified. Mail servers may hold logs. Often a fake email or document is part of an offense and proving where a faked email was sent from becomes quite relevant, and yes, ends up as evidence in court.
You may hear a prosecutor say "and on the 30th of june 2011 did you send the following email..blah..blah"There is a reason for that- they are establishing the possibility that it is actually evidence. Its not clear cut, otherwise we wouldn't have courts. You proffer evidence and convince people of it's weight. And that will continue to be the case.
Also the perpetrator is going to have to spoof my exact number to trick friends & relatives. Who I would hope would after speaking to a fake me would then text me soon or a bit after talking with comments & questions.
Demonstration from the author of the Buzzfeed article: https://soundcloud.com/cwarzel/2018-02-06t19-53-39769z-1
Demonstration using Trump's voice: https://twitter.com/LyrebirdAi/status/904595052521025536
...used AI to clone victims voice and used deep fake to social engineer customer support staff into compromising security swapping SIM card, draining bank account, opening new credit accounts and mortgages, and to call random people @55h@ts, then snicker and profit..