The first is that claiming this methodology outperforms humans is unfair; there's no expectation that Turkers are particularly good at identifying sexuality from photographs, they haven't been able to train in a remotely analogous way, and there's no comparison to "experts" in orientation recognition (if those exist).
The second is that the authors of this study use the differences extracted by VGG-Face to claim support for PHT, despite not applying an appropriate statistical test to the differences in the features between groups. This is the bigger scientific misstep in my mind, the willingness to make a strong, apparently controversial claim (I'm not too familiar with PHT or its history) without properly validating it.