It's nowhere mediocre.
It's toes-to-toes with GLM-5.3 which is one of the best Open Weight model available (With Kimi K3) for general reasoning.
I just runned it on code reviews right now and it was able to catch some thread safety issue than DeepSeek-4.1 didn't. And DeepSeek-4.1 is by no means a bad model.
That said, it hardly seems like the innovative underdog some comments here make it out to be.
GLM-5.3 and Kimi K3 released 2 and 3 months ago respectively.
For all I know anyone could train a model of similar performance by just applying already published research to a new large training run.
It's definitively not a revolution, but it is a major jump forward compared to Mistral Medium 3.5.
Worth to notice that they explicitly said they still refining the model, so we might expect something a little bit better at the end of the month
For something trained on a limited budget and 4000 GPUs, it is quite respectable.
Or if it's more a situation where I could have Opus 5.5 go through recent papers from the Chinese labs to apply these to a new large training run, resulting in a model that performs similarly to Mistral Large 4.
I honestly don't know.