It would be nice to have comparisons to Claude 3.5 for the coder model, only comparing to open source models isn’t super helpful because I would want to compare to the model I’m currently using for development work.
[1]: https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5/Qwen...
The Qwen2.5-72B model seems to do pretty well on coding benchmarks, though — although no word about Aider yet.
https://gist.github.com/victorb/7749e76f7c27674f3ae36d791e20...
AFAIK, there isn't any (micro)benchmark comparisons out yet.
Remarkable that it is at all comparable to Sonnet 3.5
Ctrl F - Code Reasoning: