Also note that with Claude models, Copilot might allocate a different number of thinking tokens compared to Claude Code.
Things may have changed now compared to when I tried it out, these tools are in constant flux. In general I've found that harnesses created by the model providers (OpenAI/Codex CLI, Anthropic/Claude Code, Google/Gemini CLI) tend to be better than generalist harnesses (cheaper too, since you're not paying a middleman).
It's not about the model. It's about the harness
Ive done it 10s of times.
Like I asked you to do this task, then you spent time looking around and now want me to pat you on the back so you can continue?
With Copilot Microsoft has basically put the meanest leanest triple-turbo'd V8 engine in a rickety 80's soviet car.
You can kinda drive it fast in a straight line if you're careful, but you can also crash and burn really hard.