To uncover a deepfake video call, ask the caller to turn sideways
metaphysic.ai
metaphysic.ai
This is a current limitation, and an artifact of the data+method but not something that should be relied upon.
If we do some adversary modelling, we can find two ways to work around this:
1) actively generate and search for such data; perhaps expensive for small actors but not well equipped malicious ones.
2) wait for deep learning to catch up, e.g. by extending NERFs (neural radiance fields) to faces; matter of time.
Now, if your company/government is on the bleeding edge of ML-based deception, they can have such policy, and they will update it 12-18-24 months (or whenever (1) or (2) materialises). However, I don't know one organisation that doesn't have some outdated security guideline that they cling to, e.g. old school password rules and rotations.
Will "turning sideways to spot a deepfake" be a valid test in 5 years? Prolly no, so don't base your secops around this.
… yes, because that worked well.
Just to be clear, Mission Impossible is not a documentary.
After all, if the artist can imagine and build a story around it, there'll be an engineer somewhere who'll go "Ah, what the hell, I could do that."
*By Golblum/Finnagle's Law, it is guaranteed said Engineer will not contemplate whether they should before implementing it, and distributing it to the world.
This another example of why we can't have nice things.
One thing I notice with colorized movies is the color of the actor's teeth tends to flicker between grey and ivory. I wonder if there are similar artifacts with deep fakes.
The extreme reaction and copypastas like this probably lead to microsoft scrapping that idea a few years later.
I hope they delete the data immediately after use.
In the eventuality that robust deepfake technology to provide fluid real-time animation of my head from limited data sources exists and someone actually wants to use it against me, they can probably find video content involving the side of my head from some freely available social network anyway.
At least with housing they don't ask me to input the information I've already sent them into their crappy website.
Bug closed, no longer an issue, overcome by events.
And, if you are willing to share, what country/bank?
I'd be more worried about people hacking into networked security camera DVRs at stores and cafes and extracting image data from there. Multiple angles. Movement. Some are very high resolution these days. Sometimes they're mounted right on the POS, in your face. Sometimes they're actually in the top bezel of the beverage coolers.
Banks are the hardest way to get this data, not the easiest one.
Is this statement based on data or a hunch? A quick google turns up a lot of bank data breaches.
Because banks have to report data breaches. Do you think every neighborhood Gas-N-Blow is publicizing, or even knows, that it's been hacked?
Mandatory reporting certainly helps IMO. Reporting should be mandatory for anyone handling PII.
(That's what Chase's fraud department tells you to do.. no joke)
Fuck these biometeric data farmers.
What do you do in this field?
What's the direction of travel on it?
What makes it worth pursuing at a commercial level? In other words - how is this tech going to be abused/monetized?
The thing with any AI/ML tech is that current limitations are always underplayed by proponents. Self-driving cars will come out next year, every year.
I'd say that until the tech actually exists, this is a great way to detect live deepfakes. Not using the technique just because maybe sometime in the future it won't work isn't very sound.
For an extreme opponent you may need additional steps. So this sideways trick probably isn't enough for CIA or whatnot, but that's about as fringe as you can get and very little generic advice applies anyway.
Or put another way, humans can't be ready and wary, constantly and indefinitely. At some point, fatigue sets in. People move in and out of the organization. Periodic reviews of security practices don't always catch everything. Why something was implemented was forgotten by institutional memory. And then there's the cost for retraining people.
Also, those that are actively using mitigations that are going to be outdated at some point are probably far more likely to be aware of how close they are to being outdated by encountering more ambiguous cases, as seeing the state of the art progress right in front of them.
As for people sticking to outdated security practices? That's a problem of people and organizations being introspective and examining themselves, and is not linked to any one thing. We all have that problem to a lesser or greater degree in all aspects of what we do, so either you have systems in place to mitigate it or you don't.
To use a Go (the game, not the language) metaphor, skilled players always assess the whole board rather than automatically make a local move in response to a local threat. What's right for one organization is not going to be right for another. Asking the caller to turn sideways to protect against deepfakes should be considered within the organization's own framework, along with the various risks involved with deepfakes, and many other risks aside from deep fake video calls.
If there is no such framework, this is no different than yoloing lines of code in a production app by a team that does not have at least some grasp of the architectural principles and constraints at play. Or worse, not understanding the “job to be done” and building the wrong product and solving for the wrong problem.
Venn diagram of people who someone wants to trick by this particular tech, those who read any security guidelines and those worthy of applying this kind of approach to in the first place is however pretty narrow for the foreseeable future. It's more of a narrative framing device to talk about 'what to do to uncover deepfake video call' as a way to present interesting current tech limitations - not that I particularly mind it.
Getting a model to work with images turned sideways is a few lines of code (just turn image sideways at training time).
Instead of pictures of faces, now they're just vertical lines.
"Come out" could mean different things in different contexts. Deepfake defence context is analogous to something like: there are cars on public roads with no driver at the wheel. And this is already true in multiple places in the world.
Some examples:
Ford - Level 4 vehicle in 2021, no gas pedal, no steering wheel, and the passenger will never need to take control of the vehicle in a predefined area.
Honda - production vehicles with automated driving capabilities on highways sometime around 2020
Toyta - Self-driving on the highway by 2020
Renault-Nissan - 2020 for the autonomous car in urban conditions, probably 2025 for the driverless car
Volvo - It’s our ambition to have a car that can drive fully autonomously on the highway by 2021.
Hyundai - We are targeting for the highway in 2020 and urban driving in 2030.
Daimler - large-scale commercial production to take off between 2020 and 2025
BMW - highly and fully automated driving into series production by 2021
Tesla - End of 2017
It certainly wasn't just Tesla who was promising self-driving cars any second now. Tesla was definitely the most agressive, but failed to meet its goals just like every other manufacturer.
--
[1] https://venturebeat.com/2017/06/04/self-driving-car-timeline...
Sorry, I couldn't resist. :)
In fact, the first (practical) one was in Boston; not in Europe.
Sorry, I couldn't resist. ;)
A number of european capitals seem to have managed to do driverless high capacity underground trains. Here in the UK, we've got a number of automated trains but for union reasons they still have drivers in the cab who press go at each station.
In the US, it looks like Detroit has a self driving line, and there are a bunch of airport shuttles. Presumably you are hitting the same union issues as us?
So what made Boston’s later entry the first “practical” one?
[Edit] Or do you mean self-driving subways? Does Boston have one already? A quick Googling suggests the opposite:
https://whdh.com/news/mbta-officials-considering-self-drivin...
Then it turned out well, we actually need a lot more cameras. Now we need high res microphones. Now we need magnets embedded in the road. Now we need highly accurate GPS maps. Now we need high power LIDAR that damages other cameras on the road. Now we need....
Each little ingredient in the soup "made only with a stone." Machine learning has utterly failed to deliver on this original promise of learning to operate a vehicle like a person, with no more sensors than a person.
I am not aware of anyone except Musk making that claim. "Machine learning" as in the statements of the main researchers, certainly did not promise anything like it.
Example, we don't say a jet ski has a current speed limitation of 80 mph, we say it can go 80, but not 81. It's a simple fact. No promise that it will be faster tomorrow, because that's not what it is, it's not its future self.
It's like they're combining startup it will always be better after you invest more money with the reality of what "is" means.
if you don't worry about deepfakes, ok. But if you worry about deepfakes, you should not be reassured that this glitch is going to save you.
I'm not a proponent, just think your argument in this context doesn't work.
If all the money on self driving cars would have been put into public transport (driverless on rails is a solved issue) and pushing shared car ownership instead, we might actually get somewhere towards congestion-free cities.
It works really well in Singapore to control congestion, and also worked well in London when they adopted it afterwards.
Public transport also works quite well in many places around the world.
It also used to work really well in North America in the past. A past when the continent was much poorer. (I'm mostly talking about USA plus Canada here.)
Public transport only works when after you step off the bus or train, you can get to your destination on foot. Density is outlawed in much of the USA and Canada.
https://www.youtube.com/watch?v=MnyeRlMsTgI&t=416s starts a good section about Fake London, Ontario. At great expense, they built a new train line. But approximately no one uses it, because you can't get anywhere when leaving the stations. The video shows an example of a station where the closest other building is about 150m away. And that's just a single building. The next ones are even further.
Land use restrictions and minimum parking requirements are a major roadblock. And just throwing money at public transit directly won't solve those.
Shared car ownership is an interesting idea. Uber can be seen as one implementation of this concept. It can be done profitably, but I'm not sure it has much impact on the shape of cities?
In the grand scheme of things, there's not much money being put into self-driving cars so far. A quick Googling gives a Forbes article that suggests about 200 billion USD.
Asking the subject to hold up a mirror and move it around pushes the matte and inpainting problems to a whole nother level (though it may require automated analysis to detect the discrepancies).
I think that too might be spoofable given enough time and data. Maybe we could have complex optical trains (reflection, distortion, chromatic aberration), possibly even one that modulates in real time...this kind of just devolves into a Byzantine generals problem. Data coming from an untrusted pipe just fundamentally isn't trustable.
If it's a high-threat context I don't think live video should be relied on regardless of deep fakes. Bribing or coercing the person is always an alternative when the stakes are high.
couldn't the same thing be said about passwords, 2FA with SMS or asymmetric cryptography?
meanwhile real IDs have been easy to replicate for decades, but are still good enough for the job.
We'll just ask them to do "the Linda Blair". If they can turn their head 360 degrees, prolly a deepfake ;P
What about doing a bunch of video calls, and asking for callers to show their profile, "to guard against deepfakes?"
I think if the caller did this without objection that would be a bigger indication that it is a deep fake than the alternative. What real person is going to comply with this?
Easy for a human, difficult for ML/AI
Base everything off the work they do, not how they look. Embracing deepfakes is accepting that you don't discriminate on appearances.
Hell, everyone should systematically deepfake themselves into white males for interviews so that there is assured to be zero racial/gender bias in the interview process.
As with any interaction with more than one adversary, there is an infinite escalation and evolution with time. And similarly then something will come up then that is unaccounted for and so on, and so on.
Ask the caller to move out of the frame and then back in again.
You will see a noticable 'step' as the face that is partially in the frame suddenly gets detected as a face and the deepfake is applied.
The only way around this is to crop the input video quite heavily - by at least one face diameter, which is a lot if the user is near the camera.
I mean that sounds a lot easier than making deep fakes work well with profile data surely?
Was a fun project, but the cat-and-mouse feeling was inescapable. For those curious, look up the DARPA MediFor project. Siwei Lyu (in the article) did a bunch of work in this space. Also see Hany Farid and Shruti Agarwal. They've worked specifically with deep fake detection.
https://419eater.com/html/tope.htm
> On receipt of the form, we will require a photograph of you, or a trusted representative as proof of identity. You will have to get a NEW photograph taken, holding two symbol of ours. The two symbols we need you to hold are a loaf of BREAD and a FISH (the name of our church). This proves that the person in the photograph is genuine. Passport or other photographs will NOT be accepted.
> (...)
> As dumb as he looks, I'm not happy. I asked for the fish to be on his head AND a loaf of bread. I got neither!
Just look it up, (or go there if you feel lazy https://digg.com/2019/bring-me-to-life-gender-swap )
1. Robust to AI improvements.
2. Blocks all kinds of faking and tampering, not just deepfakes.
3. With a bit of work can securely timestamp the video such that it can become evidence useful for dispute resolution.
4. Also applies to audio.
5. Works in the static/offline scenario where you just get a video file and have to check it.
There are probably other advantages too. The way to do such things has been known about for a long time. The issue is not any missing pieces of tech but simply building a consensus amongst hardware vendors that there's actual market demand for [deep]fake-proof IO.
In reality, deepfakes have been around for some years now but have there been any reports of actual real world attacks using them? Not sure, I didn't hear of any but maybe there's been one or two. Problem is, that's not enough to sustain a market. Attacks have to become pretty common before it's worth throwing anything more than cheap heuristics at it.
I really don't think moving our trust to unknown, unnamed manufacturers of hardware in far away places is a solution.
The solution is not going to be high tech, imho. Just like we have learned a skepticism resulting from Photoshop, we'll learn a skepticism of live video or audio.
I happen to agree with the other voices here saying this is a folly game of cat and mouse, but there are near-time methods of making this harder to fool. And that might be enough for now.
You can't make it impossible, but you can make it very difficult.
My elderly uncle almost gave $10,000 to a scammer who had convinced him that his nephew was sitting in a jail and needed this money to be paid for his bail. Luckily, he reached out to me for help and I was able to confirm that his nephew was at home, not in jail.
I honestly can't imagine some of the scams that are coming, particularly to the tech-vulnerable, if we don't do SOMETHING to make real-time deepfake video harder than it now is.
Nothing has really changed with deepfake, other than the fact that for a brief period we could be sure the person we were having a video chat with was legit because the tech didn't exist to fake it.
If you care about the identity of who you are speaking to remotely, the only solution is to cryptographically verify the other end, which just requires plain old key distribution and verification. It's just not widespread enough today for videocalls because up to now, there wasn't much need for this.
This doesn't prevent adversarial impersonation (where you cannot trust the party that want themselves impersonated). E.g., if you are an employer, and interviewing via video, you cannot tell that the person authenticated and performing on video is indeed the person you're hiring. I dont think this is a problem that _should_ be solved tbh.
However, we can also make use of the models to not properly generalize and their limitations of the training process. Anything that is out of distribution (very rare occurrence in training data) will be hard for the model: - blinking (if the model has ever only seen single frames it will create rather random unusual blinking behavior - turn around (as mentioned by the author, side views are rarer in the web) - take off your glasses - slap your cheek - draw something on your cheek - take scissors and cut a piece of your hair
The last two would be especially difficult and funny (:
https://www.nytimes.com/2007/07/02/technology/02spam.html#:~...
But I don't know much about ML so I might be wrong.
- Vague
- A variation of "works on my machine"
And is therefore tedious and doesn't add anything to the discussion. It deserves the downvotes it gets, because they are for low-quality content.
In the new titles I saw an "is site s down": the information "works here", which you seem to be calling avoidable, is in fact a basic troubleshooting step to reconstruct where the issue is.
I tried using NoScript a long time ago and felt it just broke everything and I didn't want to start whitelisting every site I use just to make the web usable. uBlock Origin is good enough for blocking ads and trackers.
Main thinking was how out of control it would get, it would probably end up looking like anti-cheat systems where its a constant cat and mouse game due to growing sophistication of deep fake models.
From a job listing, circa 2024:
- Job may require occasional lifting. (No more than 20kg)
- Expected to travel up to 25% of the year.
- Proprietary access control requires users be able to do handstands and/or simple juggling. (Feats subject to change)
- EEOC employer.
I had the same thought as many on this thread: all biometric identification is basically an arms race that moves along as new ways of gathering biometrics become convenient and ways of faking them are developed. But as you say, yubikeys also have problems. At some point it will probably be a hybrid, e.g. require a known acquaintance to digitally sign a video where you appear together.
"You look great! We just need you to blink 5 times, and you're almost done!"
"Almost done! Just show us your best side and turn your head to the left like shown above."
"Of course, you only have best sides. Just turn your head to the right like displayed above, and we can continue."
"You've almost got it! Please open your mouth and show us your teeth."
"Wow, look at you go! Just one step remaining: Tilt your head to the right like shown above."
"Now, to complete your verification, hold your national ID beside your face. Make sure it does not obstruct your head! We need to be able to see your pretty face!"
(Tongue in cheek, of course. But my banking app actually uses this kind of language, even for verification stuff, and I don't like it :D)
That would have trouble passing anti-discrimination requirements: disability (no hands), medical (bandanna covering cancer treatment hair loss), religious (burka, rasta, yarmulke, sheitel), racist (cornrow).
And trouble with: dreadlocks (can’t run fingers through), bald headed guys (as mentioned by sibling comment), and people with hairdo’s (coiffure, hairspray, topknots, plaits etcetera).
> Matt Damon’s current movie output alone, likewise, has a rough combined runtime of 144 hours, most of it available in high-definition.
> By contrast, how many profile shots do you have of yourself?
From the article
Just a total AI meltdown from one simple question.
https://www.youtube.com/watch?v=CIoBSYpgYRw
For me, I call these Eliza-isms, since it reminds me of its simple formulas like "Can you tell me more about ___" that people got so much mileage out of.
Interestingly, this is a question my father would ask patients as a paramedic who was trying to assess people's consciousness. Another would be, "what day of the week is it?".
I'd say that these technologies are just like magic - they can seem to do things that defy your expectations, but oftentimes they fall apart when looked at from a different angle.
They would be expected to answer with something matching the day/time distribution that was represented in the training data they used; like the answer to various prompts of the "current president" question is dominated by Trump, Obama and a bit of Bush and Clinton, simply because those are the presidents in the training data and the more recent events simply aren't there yet - like the many models who have no idea how to interpret the word 'Covid' simply because they have been trained on pre-2020 data even if the model was built and released later.
Instead of a "deep fake" face swap an attacker could send virtual video from a fully-virtual environment using something like an nvidia Metahuman controlled by the camera array. I think that would be pretty easily detectable today but maybe less so with an emulated bad webcam and low-res video link. The models/rigging are only going to improve in the future.
The classic "Put a shoe on your head" verification route would still defeat that, at least until someone invents a very good tool to allow those types of models to spawn and manipulate props.
But please, I don't want to be pointing to a random bus outside my window to prove that I'm not a robot/deepfake...
The degrade of news article quality > the degrade of fact-checking journalism scrutiny > the degrade of written article quality > people rather watch live stream event than reading > degrade of live stream event trustworthiness because of deepfake...
What's next? Heavily scrutinised journal articles which runs check on videos with anti-deepfake AI-based algorithm?
Oh wait we've just gone through the full cycle.
This approach works until it doesn't. How long before deep fake can handle the 90 degree profile scenario? Not saying its not a valid approach but you'd have to consider the time it takes to implement these other checks and then the time we expect deep fakes to improve in this scenario
More interestingly, what exactly are them mechanics of getting a deep fake into video call? How is it possible that a what seems like a deepfake could make its way into my Zoom? Is Zoom enabling external plugins that alter video details?
https://www.dropbox.com/s/4hf9c9kg52nxal0/Screen%20Shot%2020...
Zoom isn't aware.
And you can write your own webcam drivers to use in any program
Or use existing software with virtual webcam output like OBS or ManyCam and write a plug-in for that
Our emit a network video stream and just play your video in that kind of software instead of writing a plugin
Maybe they were just using some of those beautifying filters like chinese streamers do.
Admittedly, I use it, but I have it set pretty low. My face isn't lit up very well, and without it, in my webcam, my skin ends up looking a lot rougher than it really is.
If I set it to the max, then it just looks like a blurry mess.
If you didn't do either one of those, perhaps you now know enough so that next time you will be able to give the interviewee a chance to demonstrate whether or not they're using a "Smooth over my facial blemishes because I'm uncomfortable with how my face looks and want it to look 'prettier'." filter.
Best of luck with your interviews!
Time to start investing in closed bank branches.
* customizable 3D avatars
* customizable voices
to communicate in meetings and in communities (VR Chat style). So the origin won't be associated to your avatar or your voice, but it'll be associated with your account (like in good old chat).
An audio prompt like 'Using your <right | left> hand, repeat the numbers that I am signaling. Use <a different | the same> set of fingers from what I am using'.
Interesting times.
Encrypt/sign the feed, watermark the images with a QR code containing the sig, have an app on your phone with their pubkey and display a big green check when it matches. every pure deepfake attempt is now easily dealt with. Boom.
Why use deep fakes when you can just not and get the same result.
The combination of the three can still be defeated by someone following you, stealing the card, lifting fingerprint from a glass, and spying the PIN, but that’s a lot of trouble to go through and online identity fraud will become extinct.
IDs should never be used as secrets. That's like mixing up your username and password.
If you are concerned that all methods of communications are compromised you wouldn't suddenly trust zoom if they do some silly head movement.
I do appreciate everyone on this site contributing to my knowledge of infosec. I don't work directly in the space, but I feel the contributions on this site help educate those of us not directly working the profession.
A one time password that is shown or spoken would ensure authentication.
This is a general comment, in reality you want to do a real risk analysis, i.e. understand what you want to protect against, exactly.
The benefit is you would not have to rely on issuing commands to the remote party.
www.hownormalami.eu
Someone good answers all questions for you with your face and you get hired for few months till they realize you are a con.
Then they can use it during war.
Of course, the spatial invariants of meat-suits in motion require an understanding of volumetric structure, and not just restricted depth surface meshes.
But it's not some unencodable computational enigma.
The only thing interesting about the title is the possibility of real time deepfakes for calls. If it's not realtime then 15years ago called and they want their technique back