Full threadpoormathskills·I’m surprised it topped the reasoning models for code generation and hard prompts. The style control results are also impressive: https://twitter.com/lmarena_ai/status/1896590154871210154View on HN