HNHacker News
TopNewBestAskShowJobs

lieret

29 karma · joined July 24, 2025

submissionscomments

Show HN: New Benchmark from SWE-bench team is 0% solved

programbench.com·24 pts·lieret·
3

Show HN: All the LM solutions on SWE-bench are bloated compared to humans

twitter.com·1 pts·lieret·
0

Show HN: New eval from SWE-bench team evalutes LMs based on goals not tickets

codeclash.ai·5 pts·lieret·
1

Show HN: Randomly switching between LMs at every step boosts SWE-bench score

swebench.com·5 pts·lieret·
1

GPT-5 on SWE-bench: Cost and performance deep-dive

mini-swe-agent.com·4 pts·lieret·
3

Show HN: New SWE-bench leaderboard compares LMs without fancy agent scaffolds

swebench.com·2 pts·lieret·
0

Show HN: Mini-swe-agent achieves 65% on SWE-bench in 100 lines of python

github.com·7 pts·lieret·
4