Did you notice much improvement going from Gemini 2.5 to 3? I didn't
I just think they're all struggling to provide real world improvements
I just think they're all struggling to provide real world improvements
(I only access these models via API)
can you share your experience and data for "leap forward" ?
I noticed huge improvement from Sonnet 4.5 to Opus 4.5 when it became unthrottled a couple weeks ago. I wasn't going to sign back up with Anthropic but I did. But two weeks in it's already starting to seem to be inconsistent. And when I go back to Sonnet it feels like they did something to lobotomize it.
Meanwhile I can fire up DeepSeek 3.2 or GLM 4.6 for a fraction of the cost and get almost as good as results.