I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.
Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.