Yann Lecun himself (as of NIPS2016) was pretty critical of probability-based metrics, as those have strong dependence to the choice of model (e.g. if the model is poor the log-probability is meaningless).
In GANs, the critic and the generator are trained w.r.t. each other, reaching some kind of equilibrium. A recent proposal that seems to be "ok" for evaluating GANs was proposed by https://arxiv.org/abs/1705.05263, which is to train a separate critic on the generator, for use in evaluation (the generator never sees gradient information from this critic). This evaluation critic approximates the Wasserstein distance. One could imagine actually training the independent critic on a validation set of images not seen by the training set.
https://openreview.net/forum?id=S1EfylZ0Z
https://www.ncbi.nlm.nih.gov/pubmed/28703000
http://pubs.acs.org/doi/abs/10.1021/acs.molpharmaceut.7b0034...
Said another way: If we gave a human the task of: [make a painting of a bridge] using a handful of examples of bridges as inspiration, and they did a 1-1 copy of one of them, it would be the most efficient result. However there is generally a culturally implied task of [the new painting should not be a direct replication of one of the examples].
So this "problem" with GAN's is a novelty requirement which is not explicitly built in to the generation chain.
Copying pictures is not efficient within the scope of this problem. The whole point of these algorithms is to extract (or ideally understand) essential features of some class of objects and to be able to represent an object of such class with radically smaller amounts of data that would be required for the full description.
That is the only definition of efficiency that matters here.
Human level AI is the goal (at least mine), so every time we see something unexpected or a "failure" in ML it's worth thinking about the "failure" mode when compared with how a human could hack the system.
In general in generative models, you have some "true data distribution" P and an estimator distribution Q.
The goal is to make P and Q the same, generally by minimizing some divergence between them.
The actual objective is defined as being between the actual distributions P and Q, but because we only have so many data points, we define an empirical loss that just uses the real observations from P. So if the model makes Q just memorize the samples from P, then it actually hasn't made P and Q similar, it's only minimized the empirical loss.
One practical way to get around this with GANs is to train a conditional GAN instead of an unconditional GAN, and then run the conditioned generation task on held-out samples from the validation set. Another good and perhaps more general solution is to train an inference network and to generate reconstructions on held out data points. If they look totally different, then the model is probably not very "representative".
Now, the fact that they usually aren't identical to Google's finds is moderately impressive. So is the fact that some transitions are "smooth" - being able to move from one head/eye position to another. But the difference between drawing a face and copy-pasting someone's face onto a different hair+background is very significant, and quite often the algorithm seems to be doing the latter. (And in any case, GANs are clearly not the way humans draw faces.)
It would be interesting to see someone try to do the same thing without a neural network. How far would they get on the same training set? The dataset is 30,000 pre-aligned and cropped images. (Would be nice if there was a searchable version to make sure generated versions are not identical to something in that set.)
I bet you could get pretty far with just matching and region replacement, plus some color corrections. But not one would pay you for that.
1. If you train a conditional GAN to do image inpainting (for example, left to right), it should be quite apparent the degree to which the model is copying and pasting the training set - by running the model with "given" parts from the test set.
2. I disagree that an ideal GAN could just output the training set. I think the right conceptual framework is that any generative model is trying to produce a distribution similar to the data distribution, and we try to accomplish this by using samples from the data distribution. So if the model memorizes the training set, then it isn't actually that close to the true underlying data distribution. In likelihood-based models (for example the usual generative RNN) you can test this by evaluating likelihood on a validation set.