Deepfakes, can you spot them?
detectfakes.media.mit.edu
detectfakes.media.mit.edu
The video fakes are always blurry around the mouth and gestures don't match the words.
The audio fakes are too "clean" (too defined breaks between words and no background noise) and sound like caricatures.
Text is virtually impossible because politicians say "fake" stuff all the time (e.g. speeches someone else wrote) so you can't rely on uncharacteristic language.
I guess I'm just extra skeptical.
I did much better on them once I realized that the real speakers were much more likely to look around the room, or especially down toward the podium.
But then again, I am pretty paranoid as I do think the media and web are not out there to provide me with quality; they want to achieve money and/or power and will try to manipulate me in order to do that so I would assume fake in case of doubt.
The lack of background noise alone was a pretty instant key for me. And adding realistic background noise, appropriate echo/reverb that would match (even roughly as one-on-one conversation vs podium at press briefing) the room they are speaking in aren't even things you need deepfake for and would be simple audio post-processing.
Still the study is extremely valuable at probing the current state of things! I'm super happy they did this.
The videos with spoken words were very easy to spot, and the ones that dealt with substantive content you could tell by the intonations.
The only thing that made me score higher than random on the text segments was assuming that the stereotypical extreme things for Trump and Biden to say are usually fake.
I.e. they used a bit of truth but stretched it too far.
But nothing about either person's sentences is 'incoherent'. And 'stretching the truth too far' is something entirely different. Incoherent sentences, by definition, cannot be 'stretching the truth', because they convey no meaning.
They do, unless they don't. And that's the issue, because these text snippets could be a gaffe, or they could be GPT-3.
GPT-3 often gives you similar, somewhat meaningful sounding sentences that are a semantic mess.
https://youtu.be/4MyLwAokINc https://youtu.be/sli2Vn6Iiy4
I'm not saying everything that Trump and Biden say is like this, but the existence of so many bumbling speech examples from both makes it really hard to discern what is real and what is fake.
only one of these two does it with 100% consistency, especially if there's no teleprompter, even then he flubs it.
One wrote mostly coherent thoughts in 140 characters, the other has handlers. Guess which one got cancelled?
There’s a great Just William story where he sneaks into a waxworks exhibition and pretends to be one, and a young man comes along and points out all the defects in his utterly unrealistic anatomy. It was a bit like that.
If you weren't familiar then I could see it being hard.
A politician choses their speechwriters, works with them, and takes ultimate responsibility. The process just happens to be exactly the same as it for all other tasks they do. I would argue a demonstrated ability to choose good people and judge their work is so essential for high government office, the process involving speechwriters is a far better measure of their suitability for office.
Speechwriters, at least good ones, also don't write 'generic' texts. They will study their client's speech patterns and adjust word choices, rhythm, humor, and other patterns to fit the person's style. For offices as high as the US president, the writers tend not to change, sometimes over decades.
See the chart at the top of https://www.theatlantic.com/politics/archive/2015/01/the-lan... for word choices in the State of the Union. There are clear difference between and consistencies within each presidents' speeches.
They're still not particularly convincing, especially the audio ones sound very fake.
But I think these sorts of tests are flawed. The fact that it's a test primes you to look for these details in a way you might not if these clips were shared on Twitter/Reddit/etc. The deepfakes are still pretty obviously fake, so I don't think I'd fall for one "in the wild". But I also see this style of test used for things like "can you hear the difference between these 2 audio compression formats" or "does this virtual guitar amp sound just like the real thing", and that's just not how you consume real media. For example I'd probably never know Meshuggah switched to entirely digital guitar amps years ago if it weren't for them saying it in interviews.
When deepfakes one day do get as close to the real thing as digital guitar amps have gotten to real amps, I doubt I'd notice a random r/worldnews post is faked, even if I could still reliably pass an A/B test.
The reverse side of it is if an overwhelming amount of deep fakes get called out, people will stop trusting the real ones (we have already started stepping into that reality)
Or will they?
I found the silent videos much harder, though more in the direction that I marked a lot of them as fake when they were in fact real.
I didn't get the reasoning behind the text-only examples at all. It's impossible to verify a quote just by looking at it. We know this for several centuries - that's why academia makes such a big deal of tracking sources. So what were they trying to do here?
In general, I think the performance results here might be a bit too optimistic: It's one thing to spot a fake if you are told to specifically look for them. However, if one of those were just casually embedded in some news report while the attention of the audience is somewhere completely different, I doubt we could spot them as easily. Even moreso if the fake isn't of the two most well-known faces of the western world.
Or just sign a message and serve the signed message.
Oh no... that nail did not require the blockchain hammer.
It is true that the cryptographic primitives that blockchain is built on are relevant to many other things that we can do and have done. Blockchain in fact utilizes those properties to build a distributed shared state on which the parties have consensus. That, however, is a very special use case of those primitives. If you need message authentication, digital signatures have done that for decades and you don't need an expensive associated database to accomplish that.
Amend: The real question is how to verify cropped and recompressed snippets. In academia or HN we add a "see foo et al. [1]", so probably that needs to be encoded into the data stream, e.g. "authoritative src is https://media.whitehouse.gov/permanent-url/1234[...] with hash $hash signed by $cert_hash on $data, portions are $some_encoding_here" and better be signed as well.
You are solving the problem where an entity (the whitehouse) wants you to know that a message came from them. Yes they can use blockchain to sign it, but they can also just put it on their website, invite the press pool for an announcement, and have their PR folks reiterate the message every so often anyone asks.
This is not the problem with deepfakes. Imagine that a harmfull footage of the president saying something terrible emerges and someone has to decide if it is real or fake. This someone might be a simple citizen, a journalist, an editor etc. They check the blockchain and see that it is not signed by the whitehouse. Does that mean that it is fake? Maybe, maybe not. Maybe the president just wants to surpress it. Maybe it was an off the cuff remark while they thought nobody is recording them, or maybe they realised that they made a mistake and decided to not sign the footage after all. Which means that the person judging the authenticity of the footage haven’t gained anything from checking the blockchain.
Just as an example: during the Trump presidential campaing the now infamous “grab them by the p*sy” footage emerged. Obviously it wasn’t an official endorsed message from his campaign. Would you say it was fake just because it wasn’t signed by them?
So how would this blockchain thing help exactly?
My idea of the blockchain is that it is a chain that could be a chain of trust. Since it seems it is hard to temper with whoever owns an endpoint. But like you say there are other means for that.
So the researchers can quantify how much of people's ablitity to detect deep fakes is coming from the video/audio and not the textual content and knowledge of politician's actual political position.
As for the text: you can easily have more than 50% success rate:
- with political knowledge you know some things are unlikely to be said by a given politician
- ironically, if the text is grammatically correct, it's unlikely that it was actually spoken by a politician - real transcripts have stuttering. Of course AI could be train to contain stuttering as well, or a bad AI could spew out complete garbage. I think in case of both Biden and Trump that makes it especially hard to differentiate between bad AI and them :D.
For an example that is a bit on the nose, I recently read a quote from Sebastian Haffner about German history, that went like this:
> Germany was a much happier place at the turn of the 20th century than it was in the 80s.
Which sounds odd, coming from a left-wing, liberal journalist - unless you know from context that he was talking about the 1880s, not the 1980s. He was not wishing the Kaiserreich back, he was just comparing different historical phases of it. Nevertheless, without the context, the quote would look highly strange, even if it is technically not a fake.
This is maybe the 3rd time in my life I looked at any of the recent AI generated stuff since the breakthrough in the 2010s and they were all obvious. I got 6/6 correct on my first try. Like maybe if I never saw any normal video of any sort (NTSC, HD, LCD, CRT, mpeg) in my life then I wouldn't be able to tell, because I would have no conception of what the other 500 glitches caused by common video techniques are and if this is one of them as opposed to an edit. The first nimrod's face is wobbling and glasses are 2D and move in a different direction from his face. This is what I noticed in the first fraction of a second after opening the website, before even figuring out if this is what the test is or not.
This is literally like the 90s when people were like "WOAHHH these 3D effects are so real", and doom, etc. I really hope this isn't the state of the art people keep telling me about. It wouldn't surprise me though, as all modern politics are based on make believe concepts.
There's also the fact that this stuff is ultra low resolution, and at higher resolutions, mistakes will be even more obvious.
I did certainly feel like I learned as I went, though. Seems like the deep fakes are a bit more monotonous, regardless of the format. The real humans have more embellishment and imperfections in their tone of voice, phrasing, and facial expressions.
Update:
Come to think of it, IRL text deep fakes don't even make sense. You can make up whatever quotes you like, you don't need GPT3 to do it for you.
Exactly. John Oliver pointed out a few weeks ago that the Russian media was using stock footage of other nation's military actions (Finland?) for propaganda. No need for deep fakes. The lowest effort lie can still work.
Deep fake tech is the new photoshop. It's mostly used for good (art, memes, jokes with friends, Disney films).
You can imagine a future where deep fakes are utterly indistinguishable from real life, but I think that people will learn to recalibrate themselves against the new technologies. It may even cause us to think more critically: people frequently chime in, "how is this not photoshopped?" on social media. They'll call out videos of explosions and physics as "looking fake". They'll have the same response to purported deep faked videos.
All the same, I'm glad and appreciate that academic and policy folks are taking a close look at the technology. We need frameworks, ethics, and institutional familiarity.
They make sense for advertising and misinformation purposes.
You don't have to pay a firm to drum up interest for a brand/product, theoretically, you'd be able to generate convincing conversations about the brand between bots on Reddit or Twitter, and do it at scale. Instead of paid posts on social media, you cut out the middle man and give the perception of community support for a brand/product, despite it being artificial.
The last point goes for misinformation, as well. Want to start grassroots support for an idea/policy/person/group/place? Instead of paying teams of people to post on social media about it, you can automate the campaign and scale it with deep faked text.
I don't think deep fakes are there yet for either purpose, though, and it isn't a given that they'll be good enough for them, either.
At that point its just standard propaganda. People have been doing that since roman times.
Yeah, most of the "fake" audio clips were blatantly obvious to literally everyone with a functioning brain (and hearing ability I guess) and if you thought that represented anywhere near the state of the art in faking voices you'd be far more confident than you deserve to be.
If I hadn't already heard very good audio deepfakes I would have left the survey with far less concern than I presently have around the issue. I came in expecting that the quality of the fakes would get far more convincing over time and it really didn't, but the capability definitely exists (especially around public figures with much audio available for training)
Ha it keeps on giving, note the picture on wall at https://youtu.be/9WfZuNceFDM
I got 31/32, and the only one I got wrong was a text statement that I don't really understand the point of including. None of them were remotely difficult except for one.
I'm likewise baffled by the inclusion of text samples. It has nothing to do with deep fakes. Perhaps it's a control? I just skipped them, defaulting to the 50% skepticism slider option.
Others in this thread report similar results. I wonder if Hacker News users have a preternatural aptitude for spotting inauthentic content owing to skewed technological competence, or if that's hubris -- simple primal readings of facial expressions/inflections expose the fakes. I'd bet on the latter option. Our lizard brains are keen to detect that uncanny valley. If this recording collection is any indication, duplicity is a few years down the road...
None of these techniques are out of the budget of a college film or dedicated hobbyist.
And I disagree with you about the lipreading. There's a reason ADR in films is always done with a shot of the back of the actors head.
Even when it's done with literally the same original speaker it's pretty obvious. Or have you never fooled a stranger in the alps?
Unless I misunderstood the goals of the study of course.
It’s clear that state of the art deep fake technology can do better than what’s presented here.
What if the whole point of this exercise is to instill (false) confidence in the most people that they have the ability to distinguish a deep fake from the real thing?
For the record, I don’t believe this is the case and I fully expect that subsequent iterations of this exercise will be far more difficult.
From the presented samples I'd wager that one goal of the study might be to quantify the influence of audio-only vs video-only vs text, hence text-only, video-only, audio-only, and video with subtitles.
While text limits the decision making process to preconceived notions about the person in question and prior background knowledge, video and audio may present clear cues from human perception alone.
In cases without audio, watching the lips for sync was also a reliable giveaway.
For context: I'm mostly unaware of actions of US presidents (beyond the broad sweeps based on left vs right) and I've at most listened to Trump and Biden maybe a couple of times for a total of ~1min in the last few years - just doing this exercise has probably at least doubled my exposure.
I found the text and voice mostly impossible unless the content clued me in. Video was a little easier because I know to look for teeth, video with sound was fairly easy.
Is there a generic trick for recognising faked voices without really knowing the original (similar to looking for teeth on videos)?
Maybe we are missing some tricks to pick out the fakes.
For the majority of folks here who picked out the fakes easily, do you feel if you can do the same if the videos were of Putin speaking in Russian (with subtitles), or speaking in English with a very thick accent?
The fakes have completely clean audio with similar artificial noise (real ones have crowd, or other backround noise), the fakes have no microphone or environmental artifacts (echo/reverb, plosives, others).
What made them even easier to recognize is that accents were all wrong. They were exactly as another person said, caricatures of the original voices. I actually found that they were so obvious that most could be identified before the second or two that they make you wait to click the submit button.
If it were a foreign language, I think the success rate would be lower, but the unnatural lack of noise gives them away.
I mostly didn’t pay attention to the political content as much as I could because it seemed like a few of them were supposed to be gotchas based on that. Obviously the text only ones forced my hand.
Wish they told me at the end how many I got. Probably missed 4 or 5.
[0] https://www.nytimes.com/2003/11/25/us/technological-dub-eras...
But it wasn't. It was just shitty compression artifacts making it appear so. (Or, maybe it still was, but I'm not convinced by that evidence)
But speech cadences and patterns are super tough to counterfeit.
https://www.c-span.org/video/?507982-1/president-trump-relea...
Not boasting, just reality.I've seen some pretty astounding deepfakes, these are not part of those.
So the AI is manipulating the mouth and its problem is its top lip isnt consistent so sometimes the top lip is appearing like a horizontal 'S'. Voice impersonations are too unrealistic.
Honestly, everyone going on about how obvious the fakes are seems a little silly. If the technology is here now in this form. It might already be far more impressive elsewhere, and doubtlessly will be far more impressive in the future. We should have a discussion about what that future might look like.
This is science. They're disproving a very narrow, specific hypothesis about a particular algorithm. They're not trying to disprove deepfakes in general. It's just one experiment collecting data on one setup.
what the actual fuck am i reading? no really, i can't even understand what they mean
Anyway, THANK YOU site developers for properly handling resumption of progress and not abusing the history API. Greatly appreciated and unfortunately very rare for these sorts of sites.
just cutting things out of context is enough to completely twist reality, deep fakes are just convenient shortcuts, but hardly needed.
media going back into "anonymous sources" and "people close to the matter says" isn't helping either.
as a sciety we were going in the right direction with pgp and decentralized authority chains, but apparently all that scene is pratically dead nowadays.
Video & sound: labeled all correctly as fake/real. Text: did not even try to read.
My comments:
For video, the biggest tell for me is the mouth. Just ~5 seconds of footage was enough usually. The lips shift in size and appearance or move way too little, for example.
For audio, also a few seconds was usually enough to make a decision. The intonation of both Biden and Trump is way off in the fakes if you've spent some time listening to them during debates, interviews, etc. It has hints of the real voice but to me the fakes sounded quite a bit off. The curious fact is that I haven't listened to anything said by Trump since he left office but his characteristic voice is still very clearly embedded in my head. I guess humans are quite good at memorizing and recognizing voices. Biden to me sounds drunk/high in the fakes LOL.
Overall, all of the fakes felt pretty easy to spot.
Anybody with knowledge in the field: is this the current "state of the art" in terms of deepfakes or better results can be achieved?
>Barack Obama reads the Navy Seals Copypasta (Speech Synthesis) https://www.youtube.com/watch?v=-_MZI2YFWgI
Or here him reading Trumps inauguration speech https://www.youtube.com/watch?v=ChzEdz7aVVs
All based on https://ai.googleblog.com/2017/12/tacotron-2-generating-huma...
Better would be to let them see 10 videos all at once, half of which are fake, and ask them to divide into a fake set of 5 and a real set of 5, after looking at all of them as many times as they like. Asking "fake or real" when there is no basis to tell whether flaws should be taken as indicating "fake" or just attributed to compression seems meaningless.
Or tell people what aspects of fakeness they're trying to assess - eg, forget about video artifacts, just pay attention to the audio.
Using clips of Trump and Biden is also a bad idea. They ask you to say if you've seen one before, but aren't many people going to have seen one, but not clearly remember that, and then be influenced to think it's real by sub-conscious recognition?
Why not present pairs of videos of the same non-famous person, one fake one real, both presented with the same amount of compression, and ask one to chose which is the real one? Using many different people, of course - why would you introduce doubt about the generality of your results by using only two people?
Of course, in practice people may be less able to recognize fakes when video quality is poor, which would be useful to know, but I think one would need to investigate that issue separately, not in combination with other reasons that fakes might or might not be recognizable.
Also, the fake ones with audio or text from Trump are way too easy for anyone who has seen Trump over the years to guess (as opposed to those who know Trump from "comedy news shows" from the last few years)... because Trump wouldn't say some of those things (anti-gay marriage stuff, for example).
Biden is a bit more difficult to guess like that because he could have said anything which he thought would be popular at a given point throughout his career... but then the technology is too bad to actually be convincing enough.