I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing
I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing
One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want.
If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol.
With that said, I think it is very hard to compare models for your own use case. I do suspect there is a shiny new toy bias with all this too.
Poor Sonnet 3.5. I have neglected it so much lately I actually don't know if I have a subscription or not right now.
I do expect an Anthropic reasoning model though to blow everything else away.
It’s an amazing model but was so much faster before the hype
The servers being constantly down is the only reason I haven’t cancelled my ChatGPT subscription
My experience is as follows:
- "Reason" toggle just got enabled for me as a free tier user of ChatGPT's webchat. Apparently this is o3-mini - I have Copilot Pro (offered to me for free), which apparently has o1 too (as well as Sonnet, etc.)
From my experience DeepSeek R1 (webchat) is more expressive, more creative and its writing style is leagues better than OpenAI's models, however it under-performs Sonnet when changing code ("code completion").
Comparison screenshots for prompt "In C++, is a reference to "const C" a "const reference to C"?": https://imgur.com/a/c-is-reference-to-const-c-const-referenc...
tl;dr keep using Claude for code and DeepSeek webchat for technical questions