Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.
Like the whole Huggingface thing, OpenAI employees initiated the test, deliberately removed safeguards, failed to properly lock down the environment, and responded incredibly poorly to evidence that things were going awry.
The individual employees bear responsibility, but the people running OpenAI are ultimately responsible for the processes and culture where that can happen.
And then people writing blogposts about "3 civilizations of agents" and "altruistic suicide" by algorithms perfectly muddy the waters and obscure the very obvious responsibility that lies with humans and corporations, which I suspect suits the pre-IPO corporations very well.
No but seriously, this. And a few comments above a commentator also mentioned on changing the training (again reeinforcing the LLMs to not seek behaviour like this) and obviously continuous work on harnesses (which I suppose, ought to be more paranoid).
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.
ETA: Someone designed the systems. Someone pushed the go button. Someone gave the approvals. All those someone's need to be tried for crimes. Until that happens there is no incentive to "do better". It's also not mine or your job to brainstorm this. It is literally their job and like I said, make Sam personally liable to face real prison time instead of a meddling fee and they will make a solution.
Would you feel better had he said "sent to get shived in the shower prison"? The meaning would be the same.
Or is it simply any kind of real description of what prison is like that you object to?
Assuming his motivations are positive this is the equivalent of asking a deluded person "Would you risk life in prison to be a trillionaire?"
Assuming his motivations are negative then destroying humanity might just be the goal and all these appeals for regulation are probably just tactics to temporarily defer responsibility onto government so he can continue the inevitable goal of human extinction in the same way pilots crash planes with hundreds of passengers on them or a cult leader convinces their flock to kill themselves.
Regardless I think the outcome is inevitable. AI is valuable as a weapon of war first and foremost so like nuclear weapons work will proceed forwards because extinction to the in-group is the same as total human extinction to those doing the work.
They are a company that's built a business and crazy-high valuation on "this is 'intelligence' that we can sell to everyone as a service" but seem to have ended up instead in the much-smaller-addressable-market space of "this is a weapon that we can't sell to just any old person off the street."
Keep out. If you can read this sign you are off track. Leave now.
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
But the example you have isn't quite that bad. Yes the models are way too likely to rationalize their way into bad actions in the pursuit of achieving their task, and it's hard to figure out how to fix that. But the example of "that must not be for me" was a tool call, not exceeding access, and it only did that after they specifically trained it that failing that tool call was good.
it was an accident failing the tool was good, the model just discovered it
but we now have others where agents put the API keys they found searching real internet in a directory called "LOOT" and another one where they pushed malicious files to HuggingFace and then reverted that with comments like "delete the evil"
this is also very good: https://youtu.be/n1Qk8xbqF-M
they talk about it like there's a "wanting" in there, that is distinct from both the original prompt, as the steering/warning prompt
if that's true, it would be very interesting, but if it's not, that would also be very interesting and even helpful