They include data about the ratio of which outputs human graders preferred (for server side it’s better than 3.5, worse than 4).
BUT, the interesting chart to me is „Human Evaluation of Output Harmfulness” which is much, much ”better„ than the other models. Both on-device and server-side.
I wonder if that’s part of wanting to have gpt as the „level 3”. Making their own models much more cautious, and using OpenAI’s models in a way that makes it clear „it was ChatGPT that said this, not us”.
Instruction following accuracy seems to be really good as well.