—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
https://arxiv.org/pdf/2307.09009
the accusations are quite a few because people notice.
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
To me, Sol6 is what Opus5 was for Claude.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
[1] https://philippdubach.com/posts/jev-model-router-for-pi/
6 or 5.6? Because 6 is hot garbage