I benched frontier AI LLM models on politics, ethics, and personality traits
blackbench.ai
blackbench.ai
Grok 4.5 scores higher on liberalism (6.26) than conservatism (5.55 and is highest on conservatism.
Mistral medium scores a 3.5 on conservatism.
Why is it that grok the highest scoring on conservatism is still higher on liberalism?
I’m curious what an LLM that scored 8 on conservatism and 3.5 on liberalism would look like.