The only real essential item here is tool calling capability is it not? So I assume they tested a strong read/write/edit tool consistency?
These kinds of models might be more useful as tools to be used by larger orchestrator models, than being the orchestrators themselves.