The meat of the report for SWEs:
SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!)
So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5.
Disappointing to say the least, but somewhat expected.
SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!)
So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5.
Disappointing to say the least, but somewhat expected.
But anyway, I think it's pretty useless to look at SWE Bench's now when other way better benchmarks exist.
yeah that was the point of introducing the Pro version
OpenAI no longer recommends SWE-Bench-Pro as a benchmark: https://openai.com/index/separating-signal-from-noise-coding...