- Why wasn't this accepted to CVPR/ECCV/one of the well-established computer vision conferences? I would love to read some of the reviewers' comments about this work before I give further judgment. (If this really is some CVPR preprint, or if it actually is peer-reviewed, I'd feel much better about this.)
- Why isn't this work listed on the official curated "LFW Results" page that Erik Learned-Miller maintains? http://vis-www.cs.umass.edu/lfw/results.html Is this work so new that Erik hasn't had time to review it yet?
- Human performance on LFW is 99.2%, which is higher than what the authors think it is. The performance drops to the (claimed) 97% when we only show humans a tight crop of the face: http://www1.cs.columbia.edu/CAVE/publications/pdfs/Kumar_ICC... They discuss this difference in a paragraph in their conclusion, but I consider it dishonest to use the lower number in the abstract and imply it in the title. In fact, I consider it misleading to put "Surpassing human performance" in the title to begin with, but that's another matter :)
- Showing good performance on one dataset (LFW) is certainly not enough to show that this "outperforms humans" in the general case. Getting a state-of-the-art result on LFW these days is like squeezing a drop of water out of a rock; in my opinion, we should turn our attention to harder datasets like GBU now that these "easier" ones are solved.
I'm not terribly familiar with Gaussian processes so I'm not sure whether the math works out, but it is a pretty uncommon thing to try in this domain. (Perhaps that's what makes this work interesting, especially since this year seems to be the "Deep Learning is Eating Everyone's Lunch" year)
I also wish they describe what final-stage classifier they use for the "GaussianFace as Feature Extractor" model. Often, that's the most important step; it's strange that they didn't compare with POOF/High-dimensional-LBP/Face++'s deep-learned features/any of the other state-of-the-art feature extractors, especially considering how much worse "GaussianFace as a binary classifier" does (93% vs 97% is a huge difference in this dataset)
Just my two cents. It definitely demands further exploration. I don't see any obvious mistakes, but I'm not sure why their approach works as well as they claim it does either.
Edit: I don't mean to start a witch hunt or anything, but if the authors have the guts to put "Human-level performance" in their title, they're just begging for the community to inspect every detail and point out all the flaws in every minutiae in their work. It's our community's hot button. It's similar to the old adage about how if you want a Linux user to help you, you have to tell them how much Linux sucks. That's where much of my skepticism comes from. The most astounding papers are often the most humble, but "humble" certainly doesn't describe this work.