Were you using the Gemini model with the Claude Code harness? Otherwise, it is not an honest comparison.
Do you think Claude Code is what makes their models operate better?
And by the same token, then what would give Gemini a fair run? Because the Gemini chat app, Stitch, and the CLI are all things I’ve used and the model can’t help itself from a) saying it’s done when it isn’t; b) going off-rails; c) ignoring strict instructions after a while.