hence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2
Which itself released yesterday? You're writing, reading, and evaluating enough software in a ~36 hour period to form, reject, and form another opinion about which model makes better _architectural_ choices?
What you're currently doing is "testing out"
Likewise I have a tool-use eval set and a browser-use eval set. I use frontier models for most interactive tasks that are not home assistant or background agents.
I can imagine someone could build evals for that but I have never done so.