Opus 4.8 will burn 10k tokens trying to answer something 100% whereas GPT-5.5 will burn 2k getting it 90% which is good enough for many things.
Some personal testing on a "help me find that restaurant" prompt https://gist.github.com/nijave/2873b8b10d8c732e46264237b0755...
I was in Cotswolds, UK a couple of months ago. For those of you who don't know, it's a rural region known for its "chocolate-box" villages and honey-colored limestone architecture. Basically, you go from village to village, most commonly via bus, taking in the sights and doing touristy stuff.
When planning the trip, my sister used ChatGPT, which helpfully (and relatively quickly) found the bus schedules and times for each hop.
Midway through the day, though, we ran into a huge problem: it turns out bus schedules are different on Sundays, and more limited. Which meant we couldn't actually go to our primary destination (the Model Village), and had to cut the trip short.
Yes, ChatGPT was quick and pleasant to use, but missed a crucial detail.
Afterwards I tried it with Opus and it did not make the same mistake.
If the central question was "what is the bus schedule on `day`" and the model screws that up, it gets a fail in my book.
Also curious if Google Maps gets the timetables correct (assuming it has them).
Semi-related, I also discovered that the default web search/fetch tools are pretty primitive and Exa MCP annihilates them. I ended up doing some comparisons with Claude Code comparing built-in server-side to Exa and to a Python MCP that used SearXNG for search and Exa was a clear winner and Python+SearXNG ended up coming out roughly the same after a few cycles of letting Claude optimize the Python code and adjust SearXNG settings. Ultimately it landed on this (making some changes to optimize returning relevant context directly in the search results so the model didn't need an additional web fetch call) https://gist.github.com/nijave/604c43e3e0fdcd60f5280d3a6b109...
You need to add the actual bus schedule to context somehow (research agent, custom tool or just dump in prompt) and even the simpler modern models will be able to do the planning.
I have an example here: https://gist.github.com/nijave/2873b8b10d8c732e46264237b0755...
Tldr; all the Claude models had identical tools and some used them efficiently and verified data while others did a crap job and hallucinated responses. Additionally, Exa MCP tools generally worked better even on older/smaller model (Llama)
If I add "Research the question extensively" to your prompt at the end I get the correct answer from Haiku and Sonnet Med on first try and I've reproduced the original prompt not returning the answer.
Unfortunately every other run now gets your gist in results.
Why trust an LLM with information like bus schedules? They fuck up things like this routinely.
I then use cheaper models like GLM for personal projects but they're noticeably much worse despite being similar in benchmarks.
I think it's not only an alignment/security tool but could perhaps be used for capabilities as well.
They also target a cost-insensitive market (corporate/coding users) compared to Google/OpenAI which support massive amounts of free users.
Not sure about that one... But I think the true secret sauce for all these models is how they reason. GPT never outputs how it thinks, which "saves on tokens" but Claude absolutely tells you how it thinks, and there's people who use how it reasons about solving problems to finetune smaller open source models, with surprisingly better output.
I have got so use to the Claude personality / style of conversation that I really can't be bothered to try these other models anymore. They need to take a huge jump but that seems to be getting harder and harder because of the jumps Anthropic makes.
This Grok version is a joke if it is not even clearing the bar now. I am just getting use to and using Fable more and more. I am also trying not to forget that this is the highly delayed old Fable model that Grok can't even beat on release. There will be a new version that expands the lead in a week or two.
It all harder and harder to judge too. I just had a prompt/response this morning that Fable finally displayed its intelligence and vowed me. That is partly because anything with even the vaguest reference to biology defaults back to Opus.
I have never liked the various nerfs Anthropic has used to balance GPU (slowing down responses, quota variance, model optimizations etc) and it definitely has burned a lot of good-will.
But it has seemed that being able to look beyond the short term pitchforks has worked quite well.
It's self-reinforcing: they've got the best coding/research model, which helps them to improve their models better than the competition so they stay ahead.
Would be nice if an insider would drop some hints so that the open-source space could make some good progress.
Same as with rich person autobiographies: even when they tell you what they think it is, they can't see the path not travelled.
Yup, there's a lot of survivorship bias in those. And humans want to attribute success to some skill somehow. You cannot just have been lucky.