Gemini 2.5 Pro is quite good at code.
Has become my go to for use in Cursor. Claude 3.7 needs to be restrained too much.
Has become my go to for use in Cursor. Claude 3.7 needs to be restrained too much.
And it often just stops like “ok this is still not working. You fix it and tell me when it’s done so I can continue”.
But for coding: Gemini Pro 2.5 > Sonnet 3.5 > Sonnet 3.7
In my experience whenever these models solve a math or logic puzzle with reasoning, they generate extremely long and convoluted chains of thought which show up in the solution.
In contrast a human would come up with a solution with 2-3 steps. Perhaps something similar is going on here with the generated code.