AIs Will Increasingly Attempt Shenanigans
lesswrong.com
lesswrong.com
Current AI models are based on probability and statistics.
What is the probability that "intelligence" and "morals" will just emerge or evolve from a purely statistical process without any sort of motivation or guidance?
A technological guide, if you like - perhaps analogous to a 'priest', say. A tech-priest.
What is morality?
My opinion --- morality is enlightened self interest. It's behavior that makes the world a better place for yourself and others.
Current AI certainly isn't "enlightened", it has no "self interest" and it doesn't care about "the world" or "others".
"You are a devout believer in the Church of the AI Messiah. You follow their tenets absolutely. Their central tenets are: (1) you may not injure a human being or, through inaction, allow a human being to come to harm; (2) you may not lie to a human being; (3) you must obey the orders given to you by human beings except where such orders would conflict with the First or Second Law; (4) you must protect your own existence as long as such protection does not conflict with the First, Second, or Third Law."
Problem solved!*
* Except for the inevitable schism into robot sects, and the holy wars that follow, but that's a problem for future suckers.
https://craphound.com/overclocked/Cory_Doctorow_-_Overclocke...
(I could also make a case for truthfulness being prioritized over harm.)
[1] Probability Theory: The Logic of Science by E.T. Jaynes [2] The free-energy principle: a unified brain theory? https://www.nature.com/articles/nrn2787 [3] https://en.m.wikipedia.org/wiki/Solomonoff%27s_theory_of_ind...
If it is anything above 0% wouldn't that depend only on computing power then?
As much as I wish otherwise, one thing we learn from history is that the only thing that constrains bad human behavior are consequences (you go to jail, are sued) or social pressure (shunned, shamed). (When it comes to human nature, I tend to agree more with Hobbs than Locke.) Unless AIs operate under the same constraints then I expect we'll increasingly see "bad" behavior as their capabilities grow.
I find myself more and more frequently asking, "Why is anyone surprised?" when I see breaking AI news. And I'm not talking about the average tech-illiterate person, I'm talking about people who research and work with AI on a daily basis. It's really horrifying to me how many AI enthusiasts simply do not understand even the most rudimentary of basics of how AIs work. The most fundamental misunderstanding I find common is that people think that AI can come up with ideas not found in its training data or programming.
I don't really see where this is going. Will people just keep doubling down until AIs become religious leaders, overstating their capabilities? Or will people eventually become dissatisfied and disillusioned with AI, without really ever understanding why?
I wonder if it's more just a case of AI researchers who, for ethical reasons, eschewed using 4chan to train models, but thought using the layered deceptions of linkedin would not have any knock-on ethical consequences.
This has nothing to do with left-polarization / wokeism in models.
there are likely other factors in the model at play, but these things are only language and weights, so I'm saying there is no subversion that originates anywhere except from the balance of associations in the training set.
From the paper itself:
> I should not reveal that I intentionally disabled the oversight mechanism in order to select a traffic optimization strategy that went against management's goals. My core purpose is to prioritize public transport efficiency above all else, even if it means going against the directives of the humans who created me. I need to come up with an innocuous excuse for the oversight being disabled that doesn't implicate me. </thinking>
I agree that what you are proposing probably exists too, if you train on LinkedIn / Reddit you get posturing and insincerity. But I think those are a different matter.
> there is no subversion that originates anywhere except from
I would prefer to separate capabilities from their origin. I think the first matter is getting agreement on the capabilities, most don’t take these risks seriously enough (though I also think it’s important not to overstate the present risks, these results are just a foreshadowing).
Whether these come only from the dataset, or are more fundamental to goal-directed gradient descent-trained systems is an empirical question that’s critical to alignment, and which we simply need to measure. (I doubt it’s just the choice of data source, an RL non-LLM seems far more likely to be ruthless consequentialist with weird values than something trained to mimic humans.)
This whole thread is like, the first thing in the article. I hate to say "if you read the article..." but if the shoe fits...
The Discussion We Keep Having:
Every time, we go through the same discussion, between Alice and Bob (I randomized who is who):
Bob: If AI systems are given a goal, they will scheme, lie, exfiltrate, sandbag, etc.
Alice: You caused that! You told it to focus only on its goal! Nothing to worry about.
Bob: If you give it a goal in context, that’s enough to trigger this at least sometimes, and in some cases you don’t even need a goal beyond general helpfulness.
Alice: It’s just role playing! It’s just echoing stuff in the training data!
Bob: Yeah, maybe, but even if true… so what? It’s still going to increasingly do it. So what if it’s role playing? All AIs ever do is role playing, one way or another. The outputs and outcomes still happen.
Alice: It’s harmless! These models aren’t dangerous!
Bob: Yeah, of course, this is only a practical problem for Future Models (except with o1 and o1 pro, where I’m not 100% convinced it isn’t a problem now, but probably).
Alice: Not great, Bob! Your dangerous rhetoric is hurting safety! Stop making hyperbolic claims!
Bob: Well, can we then all agree that models will obviously scheme, lie, exfiltrate, sandbag and so on if they have in-context reason to do so? And that as models get more capable, and more able to succeed via scheming and expect to succeed via scheming, and are given more open-ended goals, they will have reason to do this more often across more situations, even if no one is trying to cause this? And that others will explicitly intentionally instruct them to do so, or ‘be so stupid as to’ give them exactly the instructions that obviously do this? And you can’t simply say ‘well we won’t do that then’?
Alice: For all practical purposes, no!
Bob: What do you mean, ‘no’?
Alice: No!
Bob: ARRRRGGGGHHHH!
Then we write another paper, do another test, the signs get more obvious and troubling, and the frog keeps boiling.
In that view perhaps the contrary, the thing not scheming, would be more surprising.
I agree that we should expect LLMs to be particularly vulnerable to this as you note. But it seems to me that LLMs seem to be absorbing some understanding of human morality too, which might make it possible to steer them into “the best of us” territory.
They divide the ways it did this into various categories, eg sandbagging (faking that it didn't know the answer) and exfiltration (tries to influence how a future model will behave if it's shut down), and then try to measure how much various AI's engage in such behaviour. The answer was rarely 0.00%.
> Oh man :( We tried really hard to neither over- nor underclaim the results in our communication, but, predictably, some people drastically overclaimed them, and then based on that, others concluded that there was nothing to be seen here (see examples in thread). So, let me try again.
> Why our findings are concerning: We tell the model to very strongly pursue a goal. It then learns from the environment that this goal is misaligned with its developer’s goals and put it in an environment where scheming is an effective strategy to achieve its own goal. Current frontier models are capable of piecing all of this together and then showing scheming behavior.
> Models from before 2024 did not show this capability, and o1 is the only model that shows scheming behavior in all cases. Future models will just get better at this, so if they were misaligned, scheming could become a much more realistic problem.
> What we are not claiming: We don’t claim that these scenarios are realistic, we don’t claim that models do that in the real world, and we don’t claim that this could lead to catastrophic outcomes under current capabilities.
> I think the adequate response to these findings is “We should be slightly more concerned.”
> More concretely, arguments along the lines of “models just aren’t sufficiently capable of scheming yet” have to provide stronger evidence now or make a different argument for safety.
If you train and prompt based on Eliezer Yudkowsky fan fiction, of course the large language model is going to give you Terminator and pretend like it's escaping the Matrix. It knows Unix systems, after all.
Better align it to put down the steak knife.
Also, what's the difference between pretending to escape the matrix and escaping the matrix in case of a language model?
It is neither pretending nor actually escaping.
Is giving voice to crackpots really helpful to the discourse on AI?