Progressive Growing of GANs for Improved Quality, Stability, Variation [video]
youtube.com
youtube.com
Then here we are, with indistinguishable 1024x1024 recreations and trippy latent space interpolations. I know not every researcher or entrepreneur has the resources of NVIDIA to train for this many days, but let's not forget, that part needs to occur only once. It makes me wonder about the day that a GAN manages to bankrupt stock photography services.
It's not like they trained this on a GPU farm. According to the paper [1], they "trained the network on a single NVIDIA Tesla P100 GPU for 20 days".
[1] http://research.nvidia.com/sites/default/files/pubs/2017-10_...
I will believe that these devices are equal when someone shows me benchmarks proving that. Until then, I am skeptical based on past experience.
Based on this, i expect these commodity GPU servers (with 10 1080Ti cards) that cost 1/10th of the DGX-1 will be huge: https://www.servethehome.com/deeplearning11-10x-nvidia-gtx-1...
https://openreview.net/forum?id=S1EfylZ0Z
https://www.ncbi.nlm.nih.gov/pubmed/28703000
http://pubs.acs.org/doi/abs/10.1021/acs.molpharmaceut.7b0034...
Said another way: If we gave a human the task of: [make a painting of a bridge] using a handful of examples of bridges as inspiration, and they did a 1-1 copy of one of them, it would be the most efficient result. However there is generally a culturally implied task of [the new painting should not be a direct replication of one of the examples].
So this "problem" with GAN's is a novelty requirement which is not explicitly built in to the generation chain.
Copying pictures is not efficient within the scope of this problem. The whole point of these algorithms is to extract (or ideally understand) essential features of some class of objects and to be able to represent an object of such class with radically smaller amounts of data that would be required for the full description.
That is the only definition of efficiency that matters here.
Human level AI is the goal (at least mine), so every time we see something unexpected or a "failure" in ML it's worth thinking about the "failure" mode when compared with how a human could hack the system.
In general in generative models, you have some "true data distribution" P and an estimator distribution Q.
The goal is to make P and Q the same, generally by minimizing some divergence between them.
The actual objective is defined as being between the actual distributions P and Q, but because we only have so many data points, we define an empirical loss that just uses the real observations from P. So if the model makes Q just memorize the samples from P, then it actually hasn't made P and Q similar, it's only minimized the empirical loss.
One practical way to get around this with GANs is to train a conditional GAN instead of an unconditional GAN, and then run the conditioned generation task on held-out samples from the validation set. Another good and perhaps more general solution is to train an inference network and to generate reconstructions on held out data points. If they look totally different, then the model is probably not very "representative".
Yann Lecun himself (as of NIPS2016) was pretty critical of probability-based metrics, as those have strong dependence to the choice of model (e.g. if the model is poor the log-probability is meaningless).
In GANs, the critic and the generator are trained w.r.t. each other, reaching some kind of equilibrium. A recent proposal that seems to be "ok" for evaluating GANs was proposed by https://arxiv.org/abs/1705.05263, which is to train a separate critic on the generator, for use in evaluation (the generator never sees gradient information from this critic). This evaluation critic approximates the Wasserstein distance. One could imagine actually training the independent critic on a validation set of images not seen by the training set.
Now, the fact that they usually aren't identical to Google's finds is moderately impressive. So is the fact that some transitions are "smooth" - being able to move from one head/eye position to another. But the difference between drawing a face and copy-pasting someone's face onto a different hair+background is very significant, and quite often the algorithm seems to be doing the latter. (And in any case, GANs are clearly not the way humans draw faces.)
It would be interesting to see someone try to do the same thing without a neural network. How far would they get on the same training set? The dataset is 30,000 pre-aligned and cropped images. (Would be nice if there was a searchable version to make sure generated versions are not identical to something in that set.)
I bet you could get pretty far with just matching and region replacement, plus some color corrections. But not one would pay you for that.
1. If you train a conditional GAN to do image inpainting (for example, left to right), it should be quite apparent the degree to which the model is copying and pasting the training set - by running the model with "given" parts from the test set.
2. I disagree that an ideal GAN could just output the training set. I think the right conceptual framework is that any generative model is trying to produce a distribution similar to the data distribution, and we try to accomplish this by using samples from the data distribution. So if the model memorizes the training set, then it isn't actually that close to the true underlying data distribution. In likelihood-based models (for example the usual generative RNN) you can test this by evaluating likelihood on a validation set.
http://research.nvidia.com/sites/default/files/pubs/2017-10_...
Source: NVIDIA
Would be nice if I could use this to convince Facebook that some fictional image is myself, though.
Can someone explain it in layman’s terms?
Convolutional Neural Networks work similarly, but with images. Instead of giving a CNN discrete features, you'll usually just use the pixels of the image itself. Through a series of layers, the CNN is able to build features itself (traditionally things like edges, corners) and learn patterns in image data. For example, a CNN might be trained on a dataset that maps images onto labels, and learn how to label new images on its own.
This video uses Generative Adversarial Networks (GANs) to actually generate new images. In this case, you have two networks "competing" against each other. One network is a traditional CNN trying to identify is an image is "real" or computer generated, and the second network tries to generate new images to trick the first network.
We've been able to generative fairly realistic small images before (usually 64x64), but doing it this well on high-resolution (1024x1024) images is unprecedented.
This work is interesting because it is the first one that has been this successful at high resolution, so the output images are large and detailed rather than postage stamp sized.
But what if we don't have a loss function? Or we don't know it? (for example, how do we even measure "what makes a face a celebrity-like face?") In that case we can train it against another network that is itself trained to differentiate between a "real" input and a "fake" input. The new network takes an image as input, and outputs a probability that the input is real or fake. We don't know the loss function, but by alternating which batch of images this network gets (fake or real), we can tell how well it does (it should estimate the real oens are real, and the fake ones fake). By training these two networks in tandem, we can use the information from the new network (the discriminator) to tell the old network (the generator) how to generate new, better images. This way we don't really need to know the loss function in advance, between the discriminator serves as our loss function.
That is the general idea. In practice, it's fairly non-trivial to get these two networks to work together nicely... often one will get much better than the other, which prevents the other from learning.
In this particular paper, they are using a technique to expand the size of the images to much larger than you would normally be able to.
Is it illegal to own a digital brain that can think up illegal porn?
But fuck it. Downvote away!
However, I'm not sure whether there are any other applications for this specific interpolation scenario that would lead to it being developed, as the effort required to make it work is likely much higher.
I don't like the horizontal sliding transition because I'm way focused on the bizarre iterations of the various targets.
Gonna have to update our camouflage patterns again to combat computer vision...
[0]: https://en.m.wikipedia.org/wiki/Generative_adversarial_netwo...
Think about this: what is the only generation that lives in times when massive recording, storage and communication is possible, but was not yet influenced by AI? Ours. When AI will want to simulate a "natural" human society, it has only us as templates to train its humanGAN on. Even our past comments on reddit and Twitter have been used many times to train dialogue systems - we're being uploaded with each post we make.
instead of "In the future", triggers me. Do you know something we do not? I am resigned to the notion that everything is up for grabs.
> I am not sure if I am a toy simulation
We need "reality discriminants". The trippy thing is if they could exist and their output is not necessarily boolean. There would be a threshold point at which beings can exist along the simulated-real spectrum, were by they can understand the output of the "reality discriminants", yet they are not real.