They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...
No, doesn't seem like it
https://openai.com/index/separating-signal-from-noise-coding...
Great catch.