> On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.
I continue to appreciate OpenAI's attempt at some honesty here, showing that they are capable enough and have skilled engineers to a point where they can recognize that slop is hated for good reason, and that there is a real issue. Compare this to anthropic, where e.g. in the Opus 5.5 announcement[1] one of the first points on the page is
> One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
This is the kind of shit that is the very reason why I stick to OpenAI and deepseek. OpenAI is simply more honest and reasonable about their models' capabilities, while delivering models that still have solid value.
Notice how the OpenAI announcement doesn't make use of anecdotes.