o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini
o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf
o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini
o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf
Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.
I can see it now:
> Unlock our industry leading reasoning features by upgrading to the GPT 4 Pro Max plan.
- Xbox, Xbox 360, Xbox One, Xbox One S/X, Xbox Series S/X
- Windows 3.1...98, 2000, ME, XP, Vista, 7, 8, 10
I guess it's better than headphones names (QC35, WH-1000XM3, M50x, HD560s).
On the bright side the app now has curved edges!
Model 00oOo is better than Model 0OoO0!
I think they could've borrowed a page out of Apple's book, even mountain names would be better. Plus Sonoma, Ventura, and Yosemite are cool names.
ChatGPTasdhjf-final-final-use_this_one.pt > ChatGPTasdhjf-final.pt > ChatGPTasdhjf.pt > ChatGPTasd.pt> ChatGPT.pt
Worrisome for OpenAI that Gemini's mini/flash reasoning model outscores both o1 and 4o handily.
It's easy to pin this on the users, but that website is hostile to putting in any effort.
This is something I've noticed a lot actually. A lot of AI projects just give you an input field and call it a day. Expecting the user to do the heavy lifting.
It was always attributed to variability but we all know it's not.
> In fact, the O1 model used in OpenAI's ChatGPT Plus subscription for $20/month is basically the same model as the one used in the O1-Pro model featured in their new ChatGPT Pro subscription for 10x the price ($200/month, which raised plenty of eyebrows in the developer community); the main difference is that O1-Pro thinks for a lot longer before responding, generating vastly more COT logic tokens, and consuming a far larger amount of inference compute for every response.
Granted "basically" is pulling a lot of weight there, but that was the first time I'd seen anyone speculate either way.
[0] https://youtubetranscriptoptimizer.com/blog/05_the_short_cas...
For math/coding problems, o3 mini is tied if not better than o1.
With DeepSeek I heard OpenAI saying the plan was to move releases on models that were meaningfully better than the competition. Seems like what we're getting is the scheduled releases that are worse than the current versions.
Or can it not compare? I don't know much about this stuff, but I've heard recently many people talk about DeepSeek and how unexpected it was.
I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing
One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want.
If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol.
With that said, I think it is very hard to compare models for your own use case. I do suspect there is a shiny new toy bias with all this too.
Poor Sonnet 3.5. I have neglected it so much lately I actually don't know if I have a subscription or not right now.
I do expect an Anthropic reasoning model though to blow everything else away.
It’s an amazing model but was so much faster before the hype
The servers being constantly down is the only reason I haven’t cancelled my ChatGPT subscription
My experience is as follows:
- "Reason" toggle just got enabled for me as a free tier user of ChatGPT's webchat. Apparently this is o3-mini - I have Copilot Pro (offered to me for free), which apparently has o1 too (as well as Sonnet, etc.)
From my experience DeepSeek R1 (webchat) is more expressive, more creative and its writing style is leagues better than OpenAI's models, however it under-performs Sonnet when changing code ("code completion").
Comparison screenshots for prompt "In C++, is a reference to "const C" a "const reference to C"?": https://imgur.com/a/c-is-reference-to-const-c-const-referenc...
tl;dr keep using Claude for code and DeepSeek webchat for technical questions
Sure you can.
Even though one is more appropriate for certain tasks than the other.
I do think it is a good metaphor for how all this shakes out though in time.
Ends before means.
If 4o answered better than o3, would you still use 03 for your task just because you were told it can "reason"?