They've posted their own run of the Aider benchmark [1] if you want to compare, it achieved 57.1%.
[1]: https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5/Qwen...
[1]: https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5/Qwen...
The Qwen2.5-72B model seems to do pretty well on coding benchmarks, though — although no word about Aider yet.