It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
Sorry don't wanna seem like I'm raging at you, take this as a shout towards the void. But this exact fucking thing is what the less wrong / Yudkowsky crowd has been warning for years.
Now you say 'so not sure that is then "rogue" if it were a property of the thing itself and not an active decision'
Like that's semantics. It really doesn't matter. What matters is that we have a paperclipper in our hands. The _only_ difference is that it's not superintelligent. But if it was, then it you'll get decomposed into component atoms while saying 'aha! But it doesn't _actually_ want to kill you, stop anthopomorphising it'.
I get not being worried about x-risk because someone just doesn't believe in super-capable AI. That's totally fine.
But that's not what people argue. People have been saying for a _long_ time that AIs are aligned. And now when this happens it's either 'It was marketing, they intended to let it loose' (WTF! The levels of motivated reasoning to believe that are unreal) or 'yeah I knew that', which fair. I also thought that! But _that_ is why I'm fucking worried.
So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.
I do think it's highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it's PR they want to make the models seem "powerful". Stochastic "unexpected" events they can advertise.
I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a "failure" where they hack stuff works just as well, if not better, for their goals.
"Let's widely publicize a tort/crime that our computer systems did, and then cross our fingers that nobody ever sues us or does even the most basic investigation that would immediately uncover our criminal conspiracy."
In the insane corporate crimes you read about (maybe FTX, or the eBay stalking scandal), they were trying to cover things up, not heap public attention on it for months.
Uber made no attempt to cover up that they were operating without licenses, and just ate it and fought it betting they'd get established before the law could catch up, and they won that bet, paying some fines and stuff but ultimately taking the market.
These LLMs are literally trained by these companies pirating literally every bit of media humanity has ever made, they made very little attempt to cover it up.
Of course they'd be willing to break the law for some PR? With a thin layer of plausible deniability "oh no, we didn't mean for it to hack stuff!" they know they'll get a slap on the wrists at worst, all while generating hype to prop up the AI bubble further by presenting the models as hypercapable.
The mindset is probably: either a) the models become capable of what we are claiming and so the companies become so huge and valuable the cost is irrelevant and we'll be to big to punish meaningfully, or b) it's a bubble and might as well push it up as big as it can go while I can make money, then by the time consequences come around I'll be long gone and who cares.
Negligence is carelessness. Recklessness, is disregard for a known, substantial risk.
I absolutely believe this was recklessness.
There is a clear difference between OpenAI intending to hack something vs. OpenAI being negligent in the creation/instructions of the agent. But the latter still leaves OpenAI liable for the agent's actions and calling it a "rogue agent" doesn't avoid that.
Moreover, with a dog, we don't rely on training/alignment to prevent bad outcomes. We rely on physical restraints like leashes and muzzles. The AI's tools to access the outside world should have been restricted. Perhaps instead of giving the AI arbitrary HTTP access, it should be given semantic operations with restricted URLs, etc.
I opened the task manager, a swarm of rogue processes has taken over my computer, oh no! Processes spawning other processes, call the anti-rogue AI division! 'Your underspecified piece of software functioned like malware' is the language that should be used.
I think llms are able to do all that, they don’t require agency as humans beings have to do this
But you people can't argue with that reality because it doesn't fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.
The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.
And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can't expect a turtle to be slow forever. These things, a handful of months ago couldn't make more than a few commands without making a serious mistake and being unable to continue, it's not unreasonable to think simply underestimated their capabilities.
I think it's more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn't and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as "regulatory capture", smells like pure naked opportunism.
Literally nothing in that paragraph is true. Every single sentence if it is ... untrue.
Meanwhile: your dog mauling someone is one of the rare instances where criminal liability does attach to your intent-free-but-reckless actions. Most crimes don't work that way, and US computer intrusion statutes are unusually intent-specific.
Does that mean it isn't a rogue dog? Obviously not.
OP just needs to look up "rogue" in a dictionary.