Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
Even Kimi K3 & GLM 5.3 are at 60.
Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.
This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.
Not sure how much benchmarks or CoT or evals or anything else means at this point.
These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
I know for some types of ML analysis, a separate model is already used to analyze the weights.
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?
So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable.
language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.
Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers.
The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional.
The solution is to point toward external, objectively verifiable evidence.
I can point to now dozens of instances of models engaging in deception. Here's plenty: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Please point to your objectively verifiable evidence.
There's nothing intrinsically "malicious" about a task to exploit vulnerable code.
They were not instructed to deceive people, they weren't instructed to attack OAI or Huggingface. The models knew they were not instructed or allowed to do either of those things but did them anyway.
btw, the fact that OpenAI doesn't have some sort of monitor/summary for the agents that they watch I find hard to believe. There's no way this is really authentic, anyway. Even a haiku summarizer would have been like "uuuh the agents are communicating" and they would have stopped it. But I bet they saw this and decided to see what would happen.
Either that, or the average poster on HN isn't nearly as critical as I had thought.
So how are you seeing through all of that to get to The Truth that you see so clearly?
Read and learn. If you have a stronger critique, post it please.
For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.
https://artificialanalysis.ai/models/gpt-6-astra
Edit:
Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.