The inputs are self-submitted photos to a dating website. Can we be certain that how people pose, smile, look, etc, isn't influenced by their intention- to find a mate of a specific gender? What if those same people were each asked to submit a second photo when they're told to 'act gay/straight' in order to fool the algorithm?
It could be that this ML algorithm is picking up on very different cues than intended. But I wouldn't discourage the researchers over it- try it again with a better data set, one built for this purpose, and discover just what the algorithm is picking up.