SWE-Bench Verified
openai.com
openai.com
I have a feeling that the filtered samples better represent actual issues in the real world. How would "AI" even start with issues like "product sometimes runs slow" or "it crashes 1/20 times with this vague error message". I think my job is secure for the moment. I would love AI tooling to better sort and categorise issues though. In a rush to "replace" people I think AI companies are focusing on the wrong things.
By asking more questions, just like a developer would. This is already available in the chat form, but that's a separate set of skills and would require the other party to give responses. It's a useful test, but I believe not in the context of what swe-bench wants to measure.
How is the "kernS: 'kern' referenced before assignment" problem an example of a good/verified sample? The comment (if it is a comment?) is using // in python and I can't tell if it's an example of something failing, or a line from the error, or something else.
I would hope we can concentrate on benchmarking on full, well structured sentences before we go into "my first hasty SO question" territory. Which don't get me wrong, is useful if you want to write very short prompts, but... we're not great with basics yet and it can be always refined to a quality prompt by the LLM in a separate step. It feels really weird to me to include that in the benchmark.