How an A.I. ‘Cat-And-Mouse Game’ Generates Believable Fake Photos
nytimes.com
nytimes.com
In particular, I think this guy is missing a pretty significant part of his head: https://static01.nyt.com/newsgraphics/2017/12/26/ai-faces/8e...
However, give it to a good Photoshop artist, like most celebrity pictures are, and I'm sure these issues will be fixed in no time.
The great idea about GANs is that they replace one of the most hard to understand parts of a neural net - the "Loss function" - with another neural net, thus making the loss function learnable. This opens up the door for a kind of unsupervised learning that was impossible to make work before. GANs are very very important also because they are almost like reinforcement learning (actor + critic = RL, generator + discriminator = GAN), and RL is supposed to be the way to AGI.
The most famous problem of GANs is instability during training and mode collapse - which is like a student learning especially for an exam (and not in general) thus optimising for the test instead of the real thing.
I must confess I haven't worker with GANs yet, but isn't that the whole point of GANs? Student is optimising for the test while the teacher is learning how to make tests as similar to reality as possible?
If I understand correctly, the main challenge is finding a way to allow teacher and student (well, generator and adversary) to learn at a similar rate, so that one doesn't stop learning because its competitor is too advanced. Is that correct?
not quite, but youre on the right path.
think about it this way: you (the generative model) are trying to predict a unit gaussian, which is just a fancy way to say bell curve. you get +1 if you predict a number in this distribution (eg 0.1 or -0.5, which is within one standard deviation of the mean of 0); you get -1 if you predict a number thats "far" from this distribution (something like 40 - which has an infinitesimally low probability of being drawn from a unit gaussian).
mode collapse, then, is when you predict 0 all the time. yes, you are technically correct but youve failed to learn the true distribution.
obviously ive simplified this quite a bit and have anthropomorphize the model, but i hope you get the gist. otherwise, the [original paper](https://arxiv.org/abs/1406.2661) is refreshingly easy to read.
Could you expand on that? The more I read from folks like LeCunn & Chollet seem to disagree strongly. Just this week Yan posted about unsupervised modeling (with or without DL) to be the next path forward, and described RL as essentially a roundabout way of doing supervised learning.
alphago zero is the canonical example of tabula rasa machine learning.
There's plenty of RL papers using RNNs and some types of memory networks.
> I didn't get lessons because I gave up and eventually took the MOOC
Udacity still didn't get him onboard. I took DLF ND because of the tutoring they promised, did GANs as my first project to be in the queue, then graduated later still with no mentoring sessions. So you didn't miss anything by dropping out. How were Ng's new lessons? Worth taking it if I did DLF + fast.ai already?
BTW, GANs main use might be allowing almost fully unsupervised learning by extending small datasets with believable data.
I've wondered if dreams are basically this. Your brain uses its world-model-prediction subsystem to generate plausible inputs against which to train its action-generation-policy subsystem. Then, in real life, the action-generation-policy subsystem can react much more appropriately and quickly to real events.
Also, toddlers' stream-of-consciousness babbling when they first start talking. They narrate everything and more than once I've wondered if it's essentially them generating their own verbal training data. When they start talking to themselves their pronunciation, grammar etc. start improving much more rapidly.
Even the 'best' headline image fails as the eyes are not the same size and the rest of the face just looks off.
Also, probably the bigger risk is not that you'll be shown an entirely fabricated image, but rather that someone could convincingly be inserted into an existing image.
Cheap trick NYTimes. Cheap.
That's one heck of a receding hairline, meaning receding out of the plane of existence.
Supporting counter evidence will become that much more important.
But that I only makes sense if you know ahead of time that the information will be valuable, and it only proves the age if there is sufficient hash power on the network, so a “private” blockchain would not be viable.
If you generate content you would have a base to test an AI to spot actual fake content. You could use video's and pictures like these to test a learning AI to spot discrepancies, then report findings in detail.
Makes me wonder if there is a future in forensics for this type of technology.
They kept training until they created images that could reliably fool the discriminator. A more powerful discriminator would just be used to create better fakes.
At the very least it seems the output is not stable: a human has to decide when to stop the Wheel of Fortune. It looks more like a series of images taken from different training sets or parameters, for the NNs I'm used to.
Caveat: I've done a lot of ML, but not GANs specifically. Is this common? How do you solve the 'where to stop' problem if the output is so unstable?
>How do you solve the 'where to stop' problem if the output is so unstable?
Looking at the discriminator loss would be a good start for that.
But compare the images from Day 5 the end. Eye colour is changing and then changing back. As is the background. And the hair colour. The position of the parting. Whether the mouth is closed or showing teeth. Day 16 is not an intermediate point between Day 9 and Day 18.
If it runs for another couple of days, would we get another version like Day 16?
That's what I mean by instability.
What they are probably doing is showing snapshots of the same noise vector (==random seed) for various epoches. Since the mapping of noise vector ~> face is totally arbitrary, the ProGAN is free to vary it as it pleases; thus, some but not perfect stability. I saw the same thing in messing around with anime GANs: a fixed set of noise vectors would show the anime faces change eye or hair color etc.
> At the very least it seems the output is not stable: a human has to decide when to stop the Wheel of Fortune.
Yeah, you can't do principled early stopping with GANs, really, because there's no held-out set and the loss is changing. I always ran until it diverged or I became impatient, and similarly with ProGAN: they ran as long as they could (takes like a week on big GPUs). To some extent, if you're using Wasserstein losses, the discriminator loss is supposed to be meaningful as a kind of absolute distance between the true image distribution and the generator distribution so you can do early stopping like 'stop if no improvement for 3 epochs'. (This is just in the pure generative approach; if you're using GANs for a semi-supervised application, presumably you can do early stopping as usual based on whatever you have held-out.)
Today its faces that feel familiar but aren't real. Tomorrow its whole cities that feel familiar but aren't real. The cities are filled with people you swear you've seen before. Perhaps the details are tailored to you personally based on the corpus of photos you've posted online.
"Narrative Dungeon Design"
https://www.youtube.com/watch?v=36lE9tV9vm0
The most amazing AI video I have ever seen, actually. I spent hours staring into it - it works great as background for many pieces of music, you can think of it as the AI version of the burning log video.
>> “ANSWER: Sorry! This was a trick question. Both images were generated by computers.”
Not really a trick question when even if you know they’re both fake that the only way to be right (confirm you are right) is to be wrong.
the technology of fakery is rising the meet the “everything is fake news” moment
Surprise! Guess I should have considered the possibility of a trick question.
Whenever I run into this often-used tactic in papers and talks, I can’t help but feel – no, the author didn’t just convince me of their point. Instead they convinced me that they don’t value being trustworthy. Often I will just stop reading the article right then. Or if I do continue I will become unforgivingly skeptical of any claim that doesn’t provide a citation that is independently verifiable.
Use of the tactic feels particularly peculiar in an article which itself grasps towards the implications of a future in which photos and videos are no longer trustworthy, a future in which personal reputation will be more meaningful.
How can such a thing be enforced to begin with?
Are companies/labs/universities/individuals themselves the only thing standing between fair play and massive misuse of realistically generated media?
That's like a chess game. We have seen AlphaGo and other MCTS implementations take the "trying to detect the deception" into account.
By the time the image is generated, it would have already been factored in.
Ha! Google Brain organised a competition on "Adversarial Attacks and Defenses" within the NIPS 2017 conference.
Reminds me of Harry Potter learning magic attack and defence arts at Hogwarts.
Imagine automated system for danger recognition on for example airport. These kind of deception attacks could make problems with these systems. Imagine if suddenly 10,20,100 airports all around globe would recognize weapons, bombs or any other dangerous items? I can imagine panic and huge news headlines badmouthing AI.
People don't trust AI. These kind of errors could only prolong proper integration, which in many ways could enhance the way we live.
After that, it feels to me like the "realness" slowly degrades.
We can impersonate voices with neural nets. We can clone timbre and style, and this tech is being used commercially by Baidu at the very least (keyword: Deep Voice 3).