I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.
I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.
https://artificialanalysis.ai/evaluations/artificial-analysi...
I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone.
Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year).
I'm very much looking forward to seeing how Qwen 3.8 27B compares to Qwen 3.6 27B next week, for example.
And the latest DeepSeek v4 Flash has extremely impressive performance for a 304B model.
I think if it was JUST "how persistent is it at reasoning", they wouldn't have bothered to make Mythos a bigger model. It costs them more to run and probably puts a bunch of strain on how much compute they have available for other users and tasks. They have a lot of incentive to not use a larger model if they can get away with it.
In the end, maybe "cost per task" is the best measure of "intelligence"? That takes into account raw size, but also how efficient they are with tokens. Like Sonnet 5 being actually more expensive than Opus 4.8 at various various benchmark tasks, showed it was very persistent but not super smart. If you just point it at easy tasks, it is probably cheaper, but hard ones you shouldn't bother because it isn't worth it even if it eventually gets there.
Also, maybe this is too tautological to even mention but.. I have to wonder how much the labs even care and test for how well models do without reasoning turned on anymore? If almost all the training has reasoning turned on, probably have access to external tools, etc it is a bit hard to say how much it proves that they are dumb if they don't do well without it. As with all of AI, the amount of real "generalization" can be hard to suss out.
But again, I think the real measure is how much can the model actually do, and how much does it cost. If a model can do something that couldn't be done before, even with the old model trying to brute force it, I think that still counts for some sort of intelligence in a practical sense anyway.