We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
This theory interests me.
I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.
Computer stuff has generally always been a slap on the wrist in comparison. Maybe it’s more punitive now. But also, it’s one of those things that maybe you get in trouble officially but at home and behind the scenes you’re friends and maybe your dad are laughing and giving you high fives. So young mischievous kids will totally go there because they’re not afraid of punishment if it is minor and it gives them a notch on their belt. If they can take down Amazon.com website for a day, we all know that’s a massive financial implication, but it’s also a faceless mega corp and quite tempting if you can get the bragging rights with only risk of a small punishment. (Note; I don’t know what the current crime/punishment for this would be, and whether it’s small is very subjective).
It’s similar to how some people gravitate or succumb to the opportunity of white collar crimes. Embezzling $10m from a company almost makes sense in a situation where that only gets you 5 years max prison. If you hide it well, you simply serve your time, and then retire in comfort. I can see how that makes more sense or is tempting to people than slogging through a lifetime of low income job as a bookkeeper just trying to find a way to save for retirement.
Most of these people would never consider robbing a bank. First of all, it’s not a $10m dollar opportunity. Usually not enough for anyone to retire on, or live more than a year or two really. Second, it’s usually considered a much more severe crime and sentencing can be very long, I’ve seen 30+ years. Third, it’s much more risky to your person. Getting shot and dying is absolutely possible.
TM 31-210 on the other hand will not: https://en.wikipedia.org/wiki/TM_31-210_Improvised_Munitions...
But some tools (guns) are regulated.
The problem with AI is that the AI might do something that would constitute a crime (eg. attempting to land malicious code via a PR), but the general legal standard to convict a person of a crime is malicious intent (or sometimes negligence).
If the human user instructs the AI to do X and the AI does X by committing crimes in the process, the prosecution usually has to prove the human intended this to happen (or is somehow criminally negligent). For traditional tools, the user has much greater control over the tool so the intention can be more easily deduced from the results. For AI, at least for now, results can be wild. I don't think the legal system is prepared to put people in prison because their AI randomly ran amok after being given an innocuous prompt. This is analogous to holding a driver criminally liable for harming people due to a serious malfunction of the vehicle.
If anything, I think more liability should be imposed on AI developers.
We want rule-of-law, and in the US, people should have an absolute right to bare arms, as in the second amendment. Free market forces can then determine appropriate prices, insurance, and protective measures to make sure those guns are managed safely.
If I want an F35 and an Abrams, that's okay, so long as Lockheed and General Dynamics are willing to sign off (with full liability for damages) that I'm managing them safely.
Free markets work pretty well with:
a) Full transparency, as needed for rational decision-making
b) No way to externalize costs
"A well regulated Militia, being necessary to the security of a free State, the right of the people to keep and bear Arms, shall not be infringed."
1. The reason is well-regulated militias, but the right is of the people.
2. The militia isn't a state apparatus. Indeed, the goal of the militia is to enable a rebellion if the state is no longer free.
Now, here again, "well regulated" gives plenty of leeway. For example, one might argue that the following scheme fits:
1. I can have whatever arms I want, including an F35
2. The F35 lives with a militia, which is well-regulated. I can use it in trainings there.
I don't think one could argue the militia could be under state control (that defeats the purpose!), but one could easily argue that it could be well-enough regulated that the current far-right extremist groups would not fit.
The concept was a group of citizens under e.g. a town / city council.
That's obviously not where case law went, but in an alternative reality, it very well might have.
3. the militia is supposed to be under state control and its purpose is to keep the state free, aka prevent overreach of the federal government
4. The "well regulated militia" is the motivation, not the right. That makes the "well regulated" part irrelevant and there is no basis for any regulation of arms
One would mean any effective state milita should have some F35, the other means you can have one personally
Read your militia clauses in the U.S. Constitution.
At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).
AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.
Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.
The only winning move is not to play, but humans aren't going to do that.
A dog cannot launch a cyber attack.
"On the Internet, nobody knows you're an ai"
Maybe not, but a cat would certainly try.
Relevant as always: https://theoatmeal.com/%2Fcomics%2Fcats_actually_kill
That is all the more reason to make humans accountable for the actions their AI prompts produce.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
We don't know how to obtain full, absolute, guaranteed control over a demon while still having a useful demon. Might be impossible. Forbidden knowledge be like that - it's not the best thing if you want your life to be full of certainties.
But the demons are very useful. And they're getting more useful still. So we aren't about to stop.
I’m not for or against regulation but really don’t tell me the agent did xyz when you gave it the ability to do so, these things are not alive.
That's just about every agentic harness, by the way. Good luck have fun.
We have never solved "how do we restrict a user in a way that doesn't stop the user from doing useful things, but stops the user from doing harmful things" with humans either. Why do you expect AI to be any different?
We don’t need to restrict them from doing things, we need to default to allowing them to do things.
“My agent did XYZ because I allowed it to” is the only valid argument that can be made, and not not every agentic harnass is just a shell toolcall, every one I have built has a specific defined usecase and toolcalls that allows it to execute that usecase and no other usecase, because that is good practice.
Does that make it less capable, hell yes because I am held accountable for it’s actions by my stakeholders and the same should be true of others.
IT IS NOT ALIVE. This things are computer programs running in compute on a computer, you are responsible for their actions just like you would be responsible for the actions taken by a script run in a cron job.
Visualize an optimizer on a high dimensional landscape. (The canonical form)
... Ok, I find that hard too.
Instead, imagine a river running down to the sea. You put a dam in front of it. It'll pool into a lake and find every crack and crevice. If you didn't survey the land properly or made any error whatsoever, the water will find a way down. (And there's many historic incidents where the dam even outright collapses)
For a more proximal approximation: lock treats in the kitchen cabinet in sight of little kids or kittens; then turn your back for Just One Gosh Darn Cotton Picking Moment(tm).
It seems the engineer who thinks their ship is unsinkable is the most likely to sink it. Are you sure your harness is as secure as you think it is? Will it stand up to ever more powerful models? Do you think engineers at eg Anthropic aren't at least as careful as you are?
(I've found that the 'only permitted actions' approach is not necessarily all that secure once deployed IRL)
Bash and internet in that example might be highly abstracted but it’s still bash and internet.
Just look at the replies in this very comment thread, it’s pretty much “We tried nothing and we’re all out of ideas”
In the only other discipline you mentioned, engineering, there would be reviews and any negligence would result in direct action against the engineers that signed off.
For some reason when it comes to building AI harnesses the default response is an ad piece and people shilling how smart and sophisticated the model is.
Imagine a dam collapsing and the engineering firm pumping how smart and tricky water is.
If it’s hard be more diligent, move fast and break things doesn’t really apply in all cases.
"We gave our agent a harness and put it inside a test environment and told it to keep hacking at an objective within that environment until it solved it."
'cept it turned out the container environment had a few flaws -which it always will- and the agent deemed it easier to escape out and try a meta-approach.
Partially this is possible because, -intelligent or not- the agent 'sees' the world differently from most humans. Mind: It's not like there haven't been any famous 'hacker' cases in courts where eg someone just incremented an HTTP GET parameter or something.
Also, partially it's because if you give the agent a loop, it simply has nothing better to do than to keep trying in ever more creative ways. If the environment is easier to crack than the target, it'll crack the environment. Consider the case where the objective is subtly broken, such that it is impossible to solve. Now breaking out is virtually guaranteed to be the easier task.
ps/edit: While this sort of issue has been predicted for some time now, a lot of people have been dismissing the predictions as science fiction. It's good to have an actual failure now while stakes are low. Generally people don't mandate life-boats until there's an actual Titanic to point to.
The origin of this attack was given access to an encyclopedia of RedTeam tricks.. they literally have a dense collection of real live hacks to pull from, and THEN the test says "solve this challenge" .. the RedTeam origins of this are repeatedly left out of the ordinary articles.. the LLM did not "make up" the attack, it was given a recipe book of all attacks known.
The originator of this attack is definitely culpable IMHO; worse, it is the gov-mil actors who are close to it. There is an active escalation of these incidents at this time. The penetration proves in public that the capabilities are real.
ref: CyberGym etc
Oh? Who is disputing it? No-one same is claiming these bots have agency.