There are ~32.000 tags. No surveillance system is using ImageNet tags to classify people into Buddhists or Not-Buddhist. Most researchers ignore these tags and focus on a 1000 classes, and know that 32k performance is not good (and these artists have no intention of making it work at all). What they are uniquely trying with this Art Project is as much research as it is activism. Note that "mantrap" is defined in synset as "A trap for catching trespassers", and that you are bound to find weird stuff among over 30k categories (imagine what you can say with the 32% most popular words in French...).
This is a photo in question: https://memepedia.ru/wp-content/uploads/2019/09/imagenet-1.p...
This is the route the network took:
person, individual, someone, somebody, mortal, soul (6978) > female, female person (150) > woman, adult female (129) > smasher, stunner, knockout, beauty, ravisher, sweetheart, peach, lulu, looker, mantrap, dish (0)
So it was (politically) correct on the first three categories, and the last one was either a crapshoot (and she could also have gotten to the subcategory of "prostitute" > "streetwalker, street girl, hooker, hustler, floozy, floozie, slattern") or she really is posing in a common "beautiful woman"-way. (The global description for this route is "A very attractive or seductive looking woman" and often triggers for females with tilted heads and lip curls).
You can turn any faces dataset into a labeled face color dataset, so if a black person being subclassified as "negro" is problematic bias or encoded racism, then all such datasets are suspect. Noisy labeled data is the norm, not some horrible exception to be avoided at all costs.