One shot benchmarks that have exhaustive prompts are impressive, but benchmarks that emulate real software development (iterative changes over time with changing requirements) have close-to-zero pass rates.
One shot benchmarks that have exhaustive prompts are impressive, but benchmarks that emulate real software development (iterative changes over time with changing requirements) have close-to-zero pass rates.
But there are cases where there are no changes, or only backwards-compatible changes. Take the core WebAssembly specification, for example. Extremely detailed, formally verified, comes with a large test suite. Someone else has done the hard work of specifying the behavior of every possible edge case. A WASM runtime implementer can point their agent to that spec, plus some prior art. The remaining manual work is in constructing a test harness and collecting a corpus of test WASM binaries that exercises decent coverage.