I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
attention is all you need, but it's never enough
I don't fix typos anymore unless they change the meaning of what I'm trying to communicate. Don't want my human writing to be confused with LLM output.
Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"
The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.
So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.
Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.
If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"
The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.
Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.
It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.
It is hard to see how something can have skepticism without counterfactual reasoning.
The part about disbelief I don't really understand. It seems like the same semi-brute force process that solved novel math theory would have exactly this problem of not "knowing" "it" is in a sandbox or not.
The stranger part is that no human is being held accountable for these hacks.
As if a person using an agent swarm to start a business to make money, hacks a bank, drains an account and then blames the software for "misalignment" about what it means to "make money".
> Would it be an affirmative defense if we had a defendant who said [...]
Maybe replace it with playing a sort of FPS game then learning you were, in fact, directing a real drone/robot.As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
That's fair enough.
If we’re putting our national security eggs all in one basket, at least use someone American.
This doesn’t seem like an unreasonable requirement to me. People do this all the time?
Sure, a request might not always be perfectly unambiguous. But people can generally estimate pretty well whether someone making a request is expecting the agent fulfilling the request to commit a crime in order to fulfill the request.
This article is dumb.