I think the framing of this question is wrong, and is a limitation of the system that I've only really seen Sam Altman point out.
These "difficult and uncomfortable statistics" aren't accurate or well-researched but if you built an LLM on an open dataset you would find the people asking these questions have an agenda and on average will always lead to a certain kind of answer. The AI doesn't "understand" the question, it's just drawing from an incredibly large bank of "average answers".
What do you expect the AI to answer with when you give it a loaded question that only really asked on a site like "stormfront.com"? I don't think this is the only area ChatGPT will fail, but I imagine there will be a layer of prompt engineering where you can get the AI to give you what you want by phrasing your question a certain way.
To try and pass off these results as objective is just wrong given the propensity for ChatGPT to be confidently wrong.