A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a "load-bearing" constraint on the task.
None of this clumsiness would be so problematic if the model didn't have such a strong drive toward autonomy. It's much like with people: there's no shame in not understanding what you're being asked to do, provided you ask clarifying questions. There's no shame in ignorance if it's wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?