A Nixon deepfake, a 'moon disaster' speech and an information ecosystem at risk
scientificamerican.com
scientificamerican.com
Some of my favorites include Sinatra singing ABBA's Dancing Queen [1], six presidents rapping the NWA classic [2], and Milton Friedman rapping 50's P.I.M.P. [3]
[1] https://www.youtube.com/watch?v=zo_w4KGifug
Note I said "great examples of what's possible today" not "great examples of JFK recording NWA's Fuck the Police".
If you're going to try to insult someone, own it and leave off the smiley. Don't hide behind that weak passive aggressiveness.
For what its worth, I don’t think the OP was intending to be insulting; this seemed more like an inside joke designed to foster camaraderie.
Probably for the same reason I disliked the Sinatra content so much, too obviously fake.
The whole of the parent comment was that single remark I quoted, before they edited the comment.
They were intending to lob a passive aggressive insult on the lighter side. They didn't like my comment, so they fired off a shot about my comment being equivalent to a questionably human, possibly mediocre GPT3 text.
I don't mind that sort of mild insult in response, I just prefer it without the unnecessary facade of the smiley.
One of those readings is just musing about the state of all this "almost there" tech, the other is just randomly insulting someone's writing.
That does seem like pretty dangerous ambiguity. Akiselev maybe wrote a blue gold dress of comments where one reading is mean-spirited. I doubt it was intentional, but could have been clearer.
Taking 'Six U.S. Presidents read "Fuck Tha Police" by N.W.A.' I certainly agree there are long sections where the synthesis is obvious - stilted sounding, overlong pauses between words, changes in tone in the middle of a sentence, background noises / changing audio quality accompanying different words and speakers, and so on.
But there are sections where it manages to generate several words in a row without those problems. Like at 23 seconds, the section "so help your black ass? / Ya goddamn right / Well, won't you tell everybody what the fuck you gotta say?" sounds pretty realistic.
If you blame the bad sections on lack of training data for older presidents, take the good sections as proof of what's possible, and imagine the entire recording will be that good in 3-5 years, then it's an impressive demo.
This one (JFK reading the Rick & Morty copypasta) strikes me as much more convincing however:
But this was based on the text of an actual speech written for President Nixen, to be delivered if the mission failed. So isn't it likely that he practiced it, just in case? And if he did, it might have been recorded, more or less accidentally. As Reagan's quip about nuking Russia was.
And so as others note, it becomes crucial to look at the historical context.
Oh this is interesting, do you have any sources where I could read more about it?
https://en.m.wikipedia.org/wiki/We_begin_bombing_in_five_min...
Nothing makes me write a source off as a bad actor more than panic mongering that /tells you/ how to feel as opposed to leaving it to you to decide. It is not only a sure sign of a manipulator because of scared people being easier to manipulate. It may be unduly harsh, caustic, or sharp but I have absolutely lost patience for this tactic.
However critical thought is not a skill we have invested in, going by what is happening in the world, and the US in particular.
> Voice is clipped
> Chin
> head bobbing
Now I'm going to have to go and watch archival Nixon footage to see if the same thing is there or not :)
I wonder if you could use a GAN to do a post-production fix on them. Teach it what actual people's head movements look like, and then get it to stabilize the image as a second pass.
I remember when cellphone cameras first came out going "ffs, who'd want that? it's a tiny, low quality photo".
The 3 minutes of 'intro blast off' (that could have been 10 seconds) and the Flash-like landing page are really problematic.
1) Collect a corpus of speeches.
2) Use style-synthesis techniques, just like in the art-style synthesis demos.
3) Input a speech about a spaceship disaster and run it in the "style" of Nixon.
The current level of AI like GPT-3 can generate fluffy text, very good fluffy text, but still can't generate a text for this kind of speech.
I disagree. Maybe the output I gave above wouldn't quite pass for the output of one of the greatest speech writers of the time, but most speeches are rubbish and I think what you can get out of GPT3 well operated is not at all rubbish.
Nonsense.
"It is with a heavy heart that we must inform you that Apollo 11 has failed to return from the moon. We have lost contact with our crew, Neil Armstrong and Edwin Aldrin. When they left earth seven hours ago their fate was unknown but now it is certain. They died as men should, ready to die for what they believed. This loss is a tremendous loss for our nation, for their families, and for mankind but their sacrifice was not made in vain: It is a testimony to courage, determination, and human achievement that will be remembered by all who witnessed it today and for all those who come for all time so long as man walks this world or any other. For this tragedy will not mark the end of space exploration; it will strengthen our resolve to continue advancing space technology and human knowledge. We have enjoyed the fruits of their labor and shared the pride of their achievement just as we now share their isolation and grief. We will honor their memory by continuing that work and by taking that faith and that dream into a future they will not live to see but helped to build. We will remember them and we will remember what they stood for: The greatness of man and the hope for a better tomorrow."
This isn't a first shot output, I had it retry a bit and guided it a little, mostly to get it to write a longer speech instead of a short quote. I think it's much better than I could have written on my own in a minute or two.
I prompted it with some mission description and said that a speech was prepared for president Nixon that in case the astronauts were stranded.
GPT3 doesn't yet do consistently GREAT output without some guidance, partially as an artefact of the generation procedure. But with a little help it does very well.
The issue is that if you just take the most likely symbol it'll rapidly go into a loop of just copying text or other degenerate behaviour. So instead, everything uses the model by sampling it-- taking less likely choices by chance weighed by the model output. Unfortunately, that means that an unlucky draw will occasional paint it into a corner. If you see that happening you can just go back and try again and you get much better output.
If that is a fair comparison depends on what your application is... if you need to to run unsupervised, it isn't consistently great. If you just need a first draft out of it or some raw ideas to turn an hour writing task into a 5 minute one, it's great for that.
I don't think this kind of manual assistance is much of a cheat either-- a real speech writer also gets exactly this sort of help from others.
[And FWIW, I did this via the GPT3 based mode in the ai dungeon video game. ... I don't have access to the GPT3 API.]
> We have enjoyed the fruits of their labor
Or this:
> This loss is a tremendous loss for our nation, for their families, and for mankind
the official speech use an increasing order, and put "friends" as an intermediate step between "family" and "nation". You want to put "family" so your don't sound insensitive and as a blow under the belt. You want to put "nation" because you want to squeeze a few votes from the tragedy and also avoid been blamed. You want to put "mankind" because you want to hide that this was a stunt in the middle of the cold war. (The official speech says "people of the world".)
There is a small risk that GPT-3 is just retrieving a distorted version of the speech from the multiples sources that were available. Like if you or me are forced to rewrite it from memory. Let's try a different scenario, like: "Yuri Gagarin got toasted during reentry, and for some reason we decided to not hide it."
In other words: Try generating a speech about "Apollo 11 disaster" without previously existing text corpus about any space disasters.
They'd pass through the Nixon questions and then say "Alright, but what about Walter Cronkite?" My true advocacy would be to use another contemporary big three anchor from the era as the deep fake, then show the real Cronkite footage during the quiz at the end
That's the real lesson IMHO, not if you can tell when you're prompted to listen for it, but when you aren't.
If you think a fully fabricated video is the key to lying to people I have a poker game you can join.
Here's a discussion of how far you can distort video today, in a way that is way more subtle and harder to rebut than a total fabrication:
I think the choice of speech is also quite telling. They chose a target with low video and audio resolution and a lot of interference that could mask some of the imperfections of their algorithm. Despite all this, oddness induced by the deepfake process is still quite evident.
Consider MP3 files. Any audiophile will tell you that a compressed MP3 file is a piss-poor representation of the true sound. The best experience is to listen to a band live, then vinyl, then FLAC.
And yet the majority of people listen to MP3 files. They strike the right balance of file size and sound clarity for an overwhelming amount of systems, doubly so since online streaming took off. So now that people have become accustomed to the sound of MP3, they are not used to anything else. They are convinced that an MP3 file is the "true sound" of the song.
Right now deepfakes is in its infancy. Personally I think it is improving at an alarming rate. If I shared the moon video on Facebook I am willing to bet most people would think it is real. They don't notice the clipped speech, the subtle double head nods. Their minds are not as critical of video as you and I are. So to them, they are convinced it is real.
What happens when only machines can detect if something is authentic or a fake? What happens when not all site administrators scan for videos and fail to mark them as illegitimate? What happens when courts use deepfakes unknowingly as evidence to convict someone, or impeach a president, or fire a worker for sexual harassment? We have already seen what happens when "verified" twitter accounts are compromised - what if a CEO puts out a video announcing some controversial new endeavor, or admission to fraud?
There are very real concerns about this technology and I believe it will very soon become a weapon that takes misinformation and turns it into very real consequences.
Imagine deepfake used to create false alibis, 911 calls, etc... That's probably not even the tip of the iceberg in coming years.
The fact we went in KNOWING it's a deepfake gave us an unmeasurable advantage/bias. For uninformed individuals that "stumble" upon this on the web, one can only imagine how it'll playout. This thought actually terrifies me... imagine a hypothetical (and conservative) 10% overall improvement AND production price reduction to this video every 6-12 months.
@SamuelAdams: Wish this site had a messaging system. Would be interesting to have a 1 on 1 discussion on this topic with you.
Example: https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice...
Its not a matter of if, it's a matter of WHEN this is economically and technologically available to the masses. The next 5-10 years is going to be a very confusing time to be alive...
Secondly, we do legitimately need to be concerned about what deepfakes will look like in five or ten years time. We're making substantial algorithmic improvements in efficiency and quality, coupled with a multi-billion-dollar race to improve computational performance for deep learning tasks.
Deepfakes may never improve (unlikely), they may gradually improve with better algorithms and more compute power, or there may be a sudden breakthrough in either area that makes 100% convincing deepfakes commonplace. We could wait until that moment to start thinking about social and technological countermeasures, but I wouldn't recommend it. If there's anything to be learned from 2020, it's that we should be investing a lot more in preparing for low-probability/high-magnitude risks.
Maybe people who use faces more than full body language are more affected by this particular one?
Got me to imagining the video where Eisenhower confesses D-Day was a failure, the actual speech he wrote does exist.
We still can't truly fake instruments, never mind the human voice.
Is face-swapping illegal? What law is it breaking?
I'm working on voice conversion (voice to voice for TeamSpeak/Discord) now.
Hell, even text suffers from this problem. Back when the UK (briefly) hit their target of 100,000 Covid-19 tests a day, the BBC News website pushed a bullshit claim that Germany was already averaging that many a month earlier. This played well into all the existing narratives about British exceptionalism, Brexit, inferiority to Europe, and the incompetence of our Covid response. It was also trivially verifiable as untrue - the German testing numbers were up on the RKI testing website, in English, and not only were they nowhere near that a month earlier, they were still well below it when the BBC published that claim. Someone at the BBC had mistaken the number of tests German labs had the capacity to process a day for the actual number of tests, which was particularly bad since literally every other part of Germany's testing process was more of a bottleneck than lab capacity. The BBC then kept this claim up and prominently linked to the article it was in on their front page for a month after they knew it was untrue. A large proportion of the UK population probably saw it. How many people do you think spotted the error? (They've since decided that because they've memory-holed it, they don't have to append a correction.)
The ability to produce fake news artificially as well won't tip the scales even slightly, because nobody with a brain trusts the "news" anyway.