Regarding seeing the metrics that they're evaluating, I have not seen real-time charts for evals. They typically take far too long for that. Instead you will kick off an eval job and either get notified when it finishes or check in every so often to sanity check some TensorBoard. I'd expect that the anomalous token usage would appear in the results, and researchers would only dig in after the fact, and after first checking that there wasn't something wrong with the instrumentation. And if the experiment was designed to be on the order of days rather than hours, it seems quite plausible that neither researchers nor infra engineers would think anything was out of the ordinary.
In any case, we are seeing more scrutiny [2], and I expect more details will be uncovered over the next few days. If this was a stunt, it was an incredibly risky one which has already somewhat backfired given the poor impression of OpenAI's internal practices and the spotlight on open models helping HF.
[1]: https://rdi.berkeley.edu/blog/exploitgym/
[2]: https://www.reuters.com/business/its-ai-agent-spent-days-hac...