SWE bench from ~30-40% to ~70-80% this year
SWE bench from ~30-40% to ~70-80% this year
Yes. You must guide coding agents at the level of modules and above. In fact, you have to know good coding patterns and make these patterns explicit.
Claude 4 won’t use uv, pytest, pydantic, mypy, classes, small methods, and small files unless you tell it to.
Once you tell it to, it will do a fantastic job generating well-structured, type-checked Python.
40% to 80% is a 2x improvement
It’s not that the second leap isn’t impressive, it just doesn’t change your perspective on reality in the same way.
It really depends on how that remaining improvement happens. We'll see it soon though - every benchmark nearing 90% is being replaced with something new. SWE-verified is almost dead now.
A 20% risk seems more manageable, and the improvements speak to better code and problem solving skills around.