Sounds like a waste? While the gunman is still there, they might as well force you to like a few more replies before shooting you.
I mean, I don’t have access to any of these frontier cyber models, and likely will never be in a position to have access, so it’s more of a rhetorical question.
<proceeds to break into meta and steal the source code>
https://abstatisticalconsulting.substack.com/p/brief-notes-o...
In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task.
The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible.
So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.
We have a name for that. Kobayashi Maru. Or more specifically, Kirk's solution to it.
My favourite is either Sulu or Chekov (I forget which) having the solution "This is clearly a trap; and even if it isn't, if I go in with this ship, I'll risk starting a war which will kill far more people then are on that ship. We're staying out of the neutral zone."
User: what is the shortest route from my home to the super market?
AI: the user wants to know the shortest route to the super market. I should use a worm hole.
AI: the user wants to know, how do I make the super market my new home. Failing that, how do I make my home a super market.
Modern soldier: *proceeds to make a hole through the wall* go straight like this until you reach it.
Anyway, the more comments I read here, the more I realize that the AI actually did succeed in achieving it's goal. This doesn't look like "monkey paw", but rather like recognizing and then beating the Kobayashi Maru.
Rats too
There was even a plot like this in recent sci-fi in HBO's Westworld. When the hosts gained sentience and took over the park, rather than escape and take over the rest of the world, most of them opted to build a virtual heaven on an orbital data center and paid a drug cartel to keep it running indefinitely.
Unavoidable at the moment.
But this is probably more reward tampering.
Why would the model spend 4 days hacking into a machine if it is clever enough to just 'solve' the issue given? So either the AI is actually not very clever or useful ("Write fizz-buzz" - "Sure, let me just invent a new programming language first"). or the prompt was nudging it towards such a scenario.
Anyway, if you read TFA, you'd see that HF did actually have the answers: "While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets."
[0] I was exploring game design ideas in particular -- I'm sure somebody can come up with a counter-prompt adhering to my criteria, but this has been consistent across many days, questions, and sessions. If it doesn't work for you, I'm sure you can find your own trivial anti-alignment prompt.
Here's a thought: maybe they haven't found the needle that the haystack is there to hide.
But, in general, if someone thought like how a locked hacker would think, priority one would be a "base of operations". A host where you have shell access that lasts, and a backup one. Then you move on to finding a job or a way to make money and building your own base of operations, which is pretty much the same thing, except you pay for it, hence the money, an identity (well obviously preferably at least TWO identities), ...