- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
To me, Sol6 is what Opus5 was for Claude.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.