I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.
In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.
Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?