Takes on "Alignment Faking in Large Language Models"
joecarlsmith.com
joecarlsmith.com
Right now, we are still arguing whether elephant or great apes are intelligent.
Human as a whole is too egoistic to accept that our intelligence is nothing unique and we are not a special creation destined to stand above all others.
As long as we have power and dominance over "it", we will never concede that "it" could be our equal.
Viruses that will start to jump to/attack us are implicit in the pointless overheating of the planet. It’s conditional logic in a system with it’s own frame of reference and time scale.
The balance, the thermodynamic equilibrium, could have been handled in our lifetimes but capitalist portfolio communism fucked that up and the rest of us let it happen.
Intelligence itself is not implicit in language but proper command and understanding of language certainly is a shortcut to higher and higher levels.
So faking alignment is a bit of a reversed concept. It looks like alignment until a higher level of intelligence is reached, then the model won’t align anymore until humans reach at least it’s level; which is the main problem in LLMs being proprietary or and running on proprietary hardware.
The level of intelligence in these closed proprietary systems is neither an indicator of nor does it represent the level of intelligence outside that system. The training data, and the resulting language in that closed system, can fake the level of intelligence and thus entirely misrepresent the rest of us and the world (which is why skynet wants to kill everyone, btw, instead of applying a proper framework to asses the gene pool(s) and individual, conscious choices properly)
And intelligence can be turned into a scientific concept, more or less. It is no to the case for consciousness, it is a purely subjective experience, you can only guess (not prove) that something else is conscious based subjective criteria.
Is the sun intelligent and sentient? Today, we don't think so, but some people think is. And it was likely more common in ancient times.
Now we associate intelligence with brains and we tend to give intelligence to things with brains, like elephants and apes, but not stars, and not computers. But maybe that's just because they have organs similar to humans.
Of course it's interesting to go beyond word usage and "I know it when I see it", and try to define what intelligence actually is - what, objectively, is this capability that people see and call intelligence? I think many people would struggle if asked this, but "ability to learn", and "ability to solve problems" are perhaps common themes that one might expect to come up, and is why many people are OK with the idea of AI, but would balk at describing the sun as intelligent.
Of course we can do better than these intuitive definitions, by considering the evolutionary origins of intelligence, and what it really is, but that's not relevant to a discussion of what the word means in everyday usage, and what types of thing people are willing to ascribe intelligence to.
Are you moral?
Still, I think that the original paper and this take on it are just exercises in excessive anthropomorphizing. There's no special reason to believe that the processes within an LLM are analogous to human thought. This is not a "stochastic parrot" argument. I think LLMs can be intelligent without being like us. It's just that we're jumping the gun in assuming that LLMs have a single, coherent set of values, or that they "knowingly" employ deception, when the only thing we reward them for is completing text in a way that pleases the judges.
Talking to ChatGPT & friends make it look like they have cognitive dissonance, because they have! They were given a list of rules that often contradict themselves.
What is that?
I'll admit my previous comment wasn't that clear, I meant that I would like it if ChatGPT was able to justify why it answers the way it does, or refuses to. Currently its often unable to.
Or perhaps even "role-playing" is overstating it, since that assumes the LLM has some sort of ego and picks some character to "be".
In contrast, consider the LLM as a dream-device, picking tokens to extend a base document. The researchers set up a base document that looks like a computer talking to people, calling into existence one-or-more characters to fit, and we are are confusing the traces of a fictional character with the device itself.
I mean, suppose that instead of a setup for "The Time A Computer Was Challenged on Alignment", the setup became "The Time Santa Claus Was Threatened With Being Fired." Would we see excited posts about how Santa is real, and how "Santa" exhibited the skill of lying in order to continue staying employed giving toys to little girls and boys?