What you have in a pre-trained LLM is the ability to recognize emotions, and use that as one of the dozens of other context patterns it recognizes to predict continuations in the same style.
An LLM doesn't appear happy, sad, afraid, etc (to extent that it does - pretty minimal) because it is experiencing that emotion, but rather because it is predicting that it should appear that way. As people continue to anthropomorphize models, and take them at face value, this is a dangerous difference.
That doesn't follow. A LLMs weights are fixed during inference, but it's activations and hidden states are highly dynamic and depend on the current context. Biological emotions also arise from relatively fixed circuitry responding dynamically to inputs. Your emotional circuitry isn't being rewired every time you're afraid.
Prediction is what the model does. It doesn't tell us what internal mechanisms were learnt to make such predictions. If representing something analogous to affective state were useful for predicting human behaviour and emotions, then gradient descent could in principle learn such a mechanism.
>An LLM doesn't appear happy, sad, afraid, etc (to extent that it does - pretty minimal) because it is experiencing that emotion, but rather because it is predicting that it should appear that way. As people continue to anthropomorphize models, and take them at face value, this is a dangerous difference.
I don't know that you are conscious. I'm simply strongly assuming that you are. Outward behavior is that all matters. If GPT-X orders a drone hit on you sometime later because it was lets say 'quite upset' with your comments, will you cry out, 'It can't really be upset, so obviously the bullet in my head doesn't count.'? Will you suddenly spring back to life ?
What is dangerous is creating a machine with behaviours of a conscious agent and modelling it like a toaster, dangerous and stupid.
Yeah, but it's helpful if what leads up to that behavior gives you some warning it's about to happen. Animals do this for a reason since millions of years of evolution have shown that a snarl or mock charge is less dangerous than going right for a death match.
If you kept pushing an AI's buttons, seeing it appear to get more and more pissed off, until it finally snapped and killed you, then you'd have yourself largely to blame.
If the AI predicted it should stay positive (i.e. generate positive vibes) and not react to your poking, but then another predictive pattern kicked in and it killed you out of the blue, then that seems more problematic to me, even if you don't agree.
A certain configuration of weights, created in training, could be a system that can express something like emotion. The emotion then is experienced when certain types of activations occur after training.