There's a bit of a cheat that's going on here though in that the model is being given the fundamental integration operations as part of the problem. That means the model hasn't had to learn what they are. It might not have needed to be given them, but it does feel like that's giving the model a leg up in the benchmarks that it wouldn't otherwise have, and when there's a direct comparison to (e.g.) DeepSeek, that's an unfair advantage.