Gemini 2.5 Pro gets 64% on SWE-bench verified.
Sonnet 3.7 gets 70%
They are reporting that GPT-4.1 gets 55%.
They are reporting that GPT-4.1 gets 55%.
Andrej Karpathy famously quipped that he only trusts two LLM evals: Chatbot Arena (which has humans blindly compare and score responses), and the r/LocalLLaMA comment section.
In practice you have to evaluate the models yourself for any non-trivial task.