You arbitrarily picked 16 possible photorealistic faces out of a total solution space of what? Millions?
Wouldn't the balance of probability be on someone in the general population of humans more closely resembling one of your 16 candidate images than any of your candidates resembling the Ground Truth image?
Doesn't the problem get both better and worse as you scale your N up from 16? That is, it would be better because one of your candidates is more likely to match the Ground Truth, but it would be worse because you've also widened your net for catching false positives?
Which makes me wonder, can we train a model on only famous people, weighted by their relative famousness?
Well, yes, but why would you want to?
Wouldn't the solution be to find a 4 different axes and pick faces that represent different endpoints. With a little psychology to help us identify what features are best to inform the public of we should be able to create a collection of photos that will be more likely to result in someone identifying the suspect than either the single photo option or the photo collection option. We would still need to test to see if it is better than the single lower quality photo option.