it is barely an improvement according to their own benchmarks. not saying thats a bad thing, but not enough for anybody to notice any difference
> Windsurf reports Opus 4.1 delivers a one standard deviation improvement over Opus 4 on their junior developer benchmark, showing roughly the same performance leap as the jump from Sonnet 3.7 to Sonnet 4.
I.e. it seems we don't get much more than new training run levels of improvement anymore. Which is better than nothing, but a shame compared to the early scaling.
Instead, ideally they’d run the benchmark tests many times, and share all of the results so we could make statistical determinations.