Taking a good model like GLM5.2 and just fine tuning it on coding can decrease real world performance due to mechanics like catastrophic forgetting. There is also other interesting behaviors were training on a broad training set can improve coding performance because there is positive transfer.
There is 100% an effort to make solid coding focused models, but it is very hard to do that without including capabilities across a broad set of adjacent tasks.
Today my "coding" sessions often enough begin with real life problems, where I discuss domain or inter-domain things, ranging from business, economics, psychology, etc. Being able to do all of that with one model is something I am willing to pay a premium for.
Of course not having to pay the premium, because the routing is smart or whatever, would be great. I just don't want to have to think about it.
intuition is that your sessions consists of 10% of domain related reasoning, and 90% of code plumbing. Those 90% could be moved to cheap and efficient specialized and focused model.
Regardless, it's fairly obvious to me that none of what I do now will require "frontier models" for much longer. Models are getting better more quickly than my problems are getting harder.
most agentic coding app can use powerful model for planning/reasoning then use "budget" model to do ground work
I've had terrible success using budget models to do ground work. The justifications that the budget models will use and document, polluting the rest of the session, are sometimes just insane. Like making code compatible with a bug that was implemented within the same session, not handling errors due to precedence in the code it just implemented, etc. I DO have success using the heavy models with lower effort, and using budget models on relatively changes post ground work. But major planning and initial ground work, I just get absolutely slop if I use a budget model.
If you're doing web stuffs, or GUI, then the budget models seem fine.
I remember them saying a few years ago that, they didn't think it was worth specializing models for code, because their general purpose models kept beating them. I guess they changed their mind? Since they did start making codex models again.