I'd be much more interested in how the biases of the models differ, and in which direction they're biased. Are there any metrics on that?
basically i present the LLM with a social situation, and ask it to take an action based on personality facets + relationship with target.
deepseek is super biased against violence. Llama 3.3 is totally okay with violence, but will never choose to "take no action", etc.