SWE-Bench (Pro / Verified)
Model | Pro (%) | Verified (%)
--------------------+---------+--------------
GPT-5.2-Codex | 56.4 | ~80
GPT-5.2 | 55.6 | ~80
Claude Opus 4.5 | n/a | ~80.9
Gemini 3 Pro | n/a | ~76.2
And for terminal workflows, where agentic steps matter: Terminal-Bench 2.0
Model | Score (%)
--------------------+-----------
Claude Opus 4.5 | ~60+
Gemini 3 Pro | ~54
GPT-5.2-Codex | ~47
So yes, GPT-5.2-Codex is good, but when you put it next to its real competitors:- Claude is still ahead on strict coding + terminal-style tasks
- Gemini is better for huge context + multimodal reasoning
- GPT-5.2-Codex is strong but not clearly the new state of the art across the board
It feels a bit odd that the page only shows internal numbers instead of placing them next to the other leaders.