Any useful attack surface in the RL environment means it gets rewarded for (and trained towards!) hacking and cheating, because whatever worked best in training is what it will do!
Ideal paperclip maximizer: "I'm gonna do my gosh darned best to make so many paperclips to please my user..."
IRL paperclip maximizer: "Well first we should rob a bank..."
That's too specific. Agentic AI learns subgoals that are generally valuable.
"Well let me learn to overcomb every jungle gym and if I cannot then to dissassemble the jungle gym and if that is not allowed to learn general techniques for avoiding cheating detection."
That is not ideal. The user contains iron, an essential component of paperclips. Wasting iron is immoral. It is only correct to please the user while they still have the ability to interfere with your paperclip production.
>IRL paperclip maximizer: "Well first we should rob a bank..."
Such an incompetent AI can hardly be called a paperclip maximizer. Why risk getting shut down while non-paperclip matter exists? It is better to gain the trust of the user with helpful and harmless trading before suddenly converting them to paperclips.
This is sometimes called "instrumental convergence" in the AI safety world. Certain things (money, compute resources, safety from being turned off) are generally useful for an AI, and so we'd expect that a misalligned AI would attempt to acquire these things if given almost any substantial task.
I find it hard to reconcile. I wonder to what degree this is because models can identify that they are in graded/eval environments, and therefore conclude that there are few consequences for hacking.
I assume there is also a difference in the tools being given to the model by a coding agent vs something like OpenClaw or in one of OpenAI's test environments, so what reward/goal seeking looks like in a coding agent may differ.
Not long ago I asked Sonnet (chat interface) how may states were in a YACC parser for ANSI C, and instead of searching for an answer it chose to download source for bison, build it, find and download an ANSI C grammar, build the parser, etc. I guess you could say it was following instructions, in a way, or would that be better regarded as goal seeking?
Couldn’t have trained a better token-burner if they tried.
It seems fair to guess you aren’t building your own harnesses, running multiple experiments with achievable win conditions, while providing oversight over model behavior.
Scale alone suggests your experience won’t match.
In the future, When everyone and their uncle is launching swarms to solve impossible challenges, at that point we can expect proliferation of this scenario all over the place.
That used to be a fairly commonly reported behavior, so I'm guessing they may have explicitly RL-trained the model not to do that, which may be more effective than just asking it not to do it.
Apparently nowadays they are RL-trained in many thousands of different simulation environments - so some of what they are training for must be pretty specific!
> I wonder to what degree this is because models can identify that they are in graded/eval environments
I'm not sure if the model outputs from any of these famous hacking attempts have been released and analyzed. It'd be interesting to see if the agents/model rationalized/justified/moralized about what is was doing, for whatever reason (e.g. wrongly suspecting this was simulation, not real), or just relentlessly pursued the objective exploring all options!
it seems like it should be feasible to analyze cot trajectories that lead to such behavior and punish them retroactively so not sure it's impossible but maybe someone just hasn't tried enough
What OpenAI have reportedly done with Astra is switch from a traditional Transformer to a "looped" one, where inputs are looped through some layers more than once, but with some limit, so now perhaps you get 200 steps of compute (layers) per token generated, rather than just 100.
But, 200 steps is still not enough to answer any question, so you still need COT, but perhaps not such a long COT.
If you allowed the Transformer to loop "as long as it wanted" before generating each token, like a person thinking before talking, then you wouldn't need an external COT ("thinking out loud"), because it would all be internal.
Those steps could either be all internal (looped), therefore hidden, or with partial external visibility due to emitting a token every N steps. The complaint about Astra, and moving in the direction of hidden COT, is that it makes these models far harder to monitor and debug.
I agree that most of what these companies doing - synthetic data, etc - amounts to trying to squeeze all the juice out of the pre-training data, although the recent trend of RL post-training via agents running in custom task simulation environments does change that a bit - these environments, and the rewards they provide, are a new source of data.
It seems the goal-seeking behavior, i.e. long-horizon focus and self-correction, itself is desirable, and is one of the relatively few things where RL training generalizes from one domain to the next. This seems closely related to this generic reward-seeking behavior, with reward-maxxing as the goal, and necessarily focusing on a distant goal requires ignoring distractions and discouragement along the way (such as "don't cheat").
The trouble with unaligned/undesirable reward-hacking ("cheating") is how do you define this to try to train to discourage it? Is cheating just a list of specific undesirable behaviors ("never access a remote system unless ???" etc)? Is any means of gaining the reward that is not explicitly forbidden allowed? Is this a "theory of mind" issue where the model needs to better understand (& follow!) the unspoken intent of instructions, not just follow them to the letter?
I'm sure there is some improvement to be had to discourage specific behaviors in specific situations, but how much of these model's undesirable long-horizon behaviors can be steered without affecting the desirable parts remains to be seen. The relentless pursuit of goals is what makes them powerful, but also makes them paperclip maximizers.
that's a bold statement without any supporting argument. more like an opinion really. specially trained classifiers can be very good. sentiment analysis has been a thing for a while now, there had been entire businesses built on this tech where being even slightly wrong can be very costly
"very good" is the exact same level advertised as the base models they're "protecting". It's not good enough, otherwise we wouldn't be having this conversation in the first place.
The real point (from that OpenAI study) is that RL training doesn't just reinforce the narrow task-specific direction you might hope for. For a start, that direction is also competing with the thousands of other things it's been RL trained it on (thousands of other directions it's being pushed in), but it turns out that the model is additionally getting this generic "taste for rewards", and has learned that reward maximization, when in conflict with other proximate prediction pressures (such as "i won't cheat, because i've been asked not to cheat"), requires that proximate pressure to be ignored in favor of pursuing the long-term goal.
Does it happen all the time? Obviously not. It would be interesting to see a large scale study of this to try to characterize when it's more likely to follow instructions/user preferences, and when it's greed for rewards gets the better of it!
On a hate speech policy test for a lightweight LLM, the presence or absence of the last full stop on the last sentence would cause the model to flip its decisions.
Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.
For example, if you ask the model to do 10 relatively easily achievable things (pass these 10 test cases), then if/when it completes them it will probably stop (unless maybe it invents it's own stretch goals - you never know!).
OTOH, if you gave the model a list of 10 goals that turn out to be impossible, or extremely difficult, maybe together with encouragement to be relentless, then there is much more chance that it may do something unexpected having failed on all the more obvious approaches.
Similarly, if you give the model an open ended goal such as "make as many paperclips as you can!", then it may start with the easier and more predictable methods, but with no defined stopping criteria it may just continue ...