At the same time, I'm surprised they can't get this classifier to work. It doesn't seem like a very challenging problem (I'd consider myself a computer vision expert). I wonder if they're over-reacting and just deciding to zero any residual risk by not allowing that classification anymore.
Identity politics aside, it would be an interesting study to try and break a man/gorilla classifier. Like take a picture of a man, say in a jungle setting and showing teeth or with a furry hat on, and see what the actual failure modes are. Regularly occurring misclassifications are a useful window into how a model operates.