Is this substantiated? Do we know this is true?
Is this substantiated? Do we know this is true?
90 minutes to build a known exploit -> much much longer to create two zero-days and escape the sandbox then hack HF == No tracking of the time it worked on that one question.
Average tokens required to complete the evaluation -> tokens required for two zero-days, network traversal, credential stealing, remote system hacking == No tracking of token usage EXPLODING at some point before it finished the whole benchmark.
Etc, etc.
Perhaps people aren’t quite understanding what it takes to discover an exploitable zero-day for your exact current system to achieve the exact goal you have right now, then do it twice.
No, you are assuming that "consuming a lot more time, tokens, other metrics" is indicative of a problem that needs to be mitigated immediately. I don't see why this would be true in the context of model evaluations. More aggressive consumption could easily mean "the model is dumb as fuck" or "the model is trying interesting things that we can learn from after the fact."
If you believe in your own containment (which obviously they did and shouldn't have) I don't see why it'd be obvious that there's something to stop. The only harm that could be done is burning tokens, which in this context might very well be synonymous with "generating experimental data."