Human: "Act like you're in pain."
Computer: "Aaugh! It hurts! Why?! No more!"
Human: "OMG! The computer feels pain!!"
Human: "Act like you're in pain."
Computer: "Aaugh! It hurts! Why?! No more!"
Human: "OMG! The computer feels pain!!"
Human: "Let me just poke this vector and see what happens"
LLM: "OUW!"
The underlying experiment didn't tell the LLM what to do. Instead, the experimenters modified a vector and observed the outcome; thus showing that there is a vector that makes an LLM go "Ouw" . People called it a "pain vector", because that's easier to remember than , idk, LVF12345.
I’ll continue what I say every time this discussion (be it consciousness, pain, self awareness, etc) comes up:
Assertions as monumental as these require equally monumental evidence. And despite certain labs and their employees making statements, their actions betray that they do not believe this to be the case.
If every LLM session were a conscious being, existing regulation for animals (controversially considered both sentient and economically useful) would need to be applied on each of these sessions. I suspect no lab will take that conclusion, for obvious reasons.
Paper says you poke the vector(s), the LLM exhibits aversive behaviours. You picks your scoring, you gets your operationalization.
That gets you an empirical result. Short of a replication failure, we can't really argue with that anymore.
What we can do is be very cautious as to how we interpret it.
People and animals have emotions because they were developed under evolutionary pressure that made emotional animals better fit to survive. An animal that can feel anger or fear is more fit to survive than one that doesn't. But LLMs aren't put under those same pressures. Their evolutionary pressure is to be a good text predictor.
But we can't subject them to the type of pain signal they experience during inference, because to the LLM, whether it's acting happy or pained, it's merely outputting what it's trained to be the most likely text to follow what came before.
Edit: s/are/aren't/
Now -while it's doing that- if you poke at certain vectors, its predictions will veer off course in interesting ways.
This shows that the activation vectors exist, and that their modification provides a causal contribution to the output.
Sure there's interesting consequences of that. But if we say that's the take-home message, that's good enough for me for one day!