10,271 karma · joined February 13, 2020
Occams razor vs. Hanlon's razor?
OpenAI: shocked pikachu how could it hack?!?
They even went to a black hat conference and somehow boasted about it.
Oh, I've used balsa wood for the nuclear core containment, it didn't work! Bad radiation, bad radiation!
US companies can meddle as much as they want in US affairs, I don't care. But not in the country I'm in.
"Power consumption is down 45% - 50% annually."
Well is it 45 or 50? Is this a measurement or prediction or a random number? From an engineer or marketing? If prediction, why not saying so? Is this a range for several years depending on elevator usage?
Things like this irk me a lot.
This was part of evaluating cyber security of their frontier models and they had a "sandbox" which, and I'm not a security researcher, looks not adequate from the first look.
I think "resilient" just means "backup copy" and I do think (IANAL) it is illegal to destroy emails when asked for them in the US.
Or was your comment ironic? Sorry, German, irony impaired.
I think that was the requirement, but yes, the cache could have been offline.
Still then they could have hacked it to create the message boards - but not use it to access the internet.
"Resilient replicas of your data will live in the US"
?
reducePrivs()
serve get(package) {
secPackage = secure(package)
getBinaryFromArtifactory(secPackage)
}
I would think the code is very small and easier to verify,
it doesn't especially have the ability to write files and act as a message board as Artifactory did.And even if the agent tries to hack that, the attack surface is 1000x smaller and the possibility also much smaller.
But I'm not a security researcher, would love to see your hack to learn something (because that is what I do to sandbox agents that need services).
Someone had to give the agent some instructions, like "hack X", "Find exploit for Y" or "Do whatever havoc you can think of" - either way the agents didn't not act on their own. They might hack HF on their own, today Claude decided to play sound through the sound pipeline I instructed it to build and measure it to see if it works, but it didn't install the sound pipeline because it hasn't had anything better to do but because I instructed it that way.
Then security researchers create a black hack talk.
$$$
I've now watched the video on the idea that your write-up was misleading.
BUT the video is much worse. For two months with highly dangerous agents agents were hacking a service and none of the researchers watched (drank coffee for 2 months, didn't say).
THEN they found the hack, removed the message board.
AND the agents found another way to create a message board, on the same service, and the researchers again - after the agents having hacked a service - do nothing - like monitoring the hacked service or tightening the sandbox.
WOW!
THEN agents hacked OpenAI infrastructure, and the researchers did nothing.
THEN the agents hacked HF.
The video does not explain why the agents run for two months unattended. They claim for model training, but don't explain how letting run agents without proper sandboxes (One might think they had written a small proxy to Artifactory with 'list packages' & 'install package <x>' to prevent leaks or hacks of the service, but no, their sandbox is no sandbox at all, but security researchers!)
But it makes a nice PR presentation on agent capbilities.
CUI BONO!
----
I just find it unbelievable that agents on their own collaborated months after an initial prompt without any guidance or direction towards a goal - which is what your write-up seems to imply with sentences like:
"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."
"discover this new informal message"
How? Why? What was their original task?
And on the researchers:
If this is highly dangerous work, why wasn't it monitored?
"Beyond a whole lot of online conspiracy theories [...]"
The agents did something 'ABC' then found the informal message board without direction, then collaborated on that months later without any guidance from humans ("like try to hack/exploit ABC").
I personally think putting people who disagree with OpenAI PR to pump the company value in a "conspiracy" box is quite a weak move.
I work with Claude Code daily for a long time now, it never started to work without a prompt or direction. It never idled and then said, "Wait, I could hack Amazon today! Oh there is a message board of other agents who already hacked a way into the internet, how convenient and quite at the right time!"
I do think strong claims need strong evidence.