or consider building your own, it's not that hard and you learn a lot about why these tools have certain peculiarities in the process
tl;dr the system prompts and context engineering matter a lot, they people building these tools have unlimited AI tokens and toy problems, so they aren't great when you try them on IRL problems
I use gemini-3, gemini-cli (agent cli tool with specific prompts (that suck), sucks)
It is also filled with childish wording, very unprofessional
Also very hard algorithmic problems. Also bugs that Claude Code or Codex CLI are completely stuck on and can't find, Gemini 3 using Gemini CLI will go in and find the impossible. It's incredible when it does this.
Claude Code burns through context like crazy. With or without Serena.
Codex CLI using GPT 5.1 Codex Max Extra High is really good and uses context more efficiently. But Gemini 3/CLI is 10X more efficient with context.
But for the past 3 days I've had to endure absolute torture and I think it's the actual Gemini 3 back-end model that is having major issues.