I distinctly remember someone making an experiment by asking ChatGPT to write jokes (?) about different groups and calculating the likelihood of it refusing, to produce a ranking. I think it was a medium article, but now I cannot find it anymore. Does anyone have a link?
EDIT: At least here is a paper aiming to predict ChatGPT prompt refusal https://arxiv.org/pdf/2306.03423 with an associated dataset https://github.com/maxwellreuter/chatgpt-refusals
EDIT2: Aha, found it! https://davidrozado.substack.com/p/openaicms An interesting graph is about 3/4 down the page, showing what ChatGPT moderation considers to be hateful.