KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
1 - https://bench.killswitch-lang.org/KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
1 - https://bench.killswitch-lang.org/- capped per-task budget and time limit
- No internet access
- different harnesses mixed
I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.
Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.
In my experience models just don't take forever to mark tasks as done.
For my coding usage I don't set time limit so benchmarks that do so are providing me less interesting use cases.
With that said, I can understand having budget limits for expensive LLMs for those of us that don't have infinite VC money.
As for no internet access, the issue is that for most problems we do want Llms to be able to search docs, API specs, GitHub issues, etc. So not allowing that usually just favours larger LLMs that were able to memorize more data, not necessarily smarter ones when both have internet access.