Briefly skimming through this paper, it appears that these numbers are not a fair comparison, as this paper uses the unrestricted protocol of LFW[1], whereas the other methods in the ROC curve shown in the paper are using the restricted protocol. As you might imagine, the latter is more restrictive -- specifically in terms of amount of training data allowed. And as I mentioned in my previous comment, training data is king in these kind of systems -- more is always better.
To go slightly out on a limb, I think more significant than the new theoretical model proposed in this paper is probably the use of lots of different types of datasets for training. (Significantly more data >> more complicated models, most of the time.) But I'd have to read the paper much more carefully to be sure about this.