> I would say that talking about a thermostat's goals is an even stronger example of anthropomorphism.
Our disconnect might be a subtle difference in what we mean when we say "goals"? The thermostat is performing actions which minimize an error and if you give the thermostat extreme amounts of power in service of that minimization then you might reach an unpleasant world-state. Nothing in that description used any analogies to human behavior. I used the word "goal" because that seems like a good description of what is happening, but if for you "goal" denotes the thing which humans do then feel free to substitute a different word.
I agree it is silly to be afraid of thermostats but that's largely because there are not any compelling reasons to give a thermostat much power or intelligence.
> What I don't follow is the scenario where the drug discovery program "wants" to show you bad drugs but shows you good ones instead because it thinks you'll eventually put it in charge of the FDA.
I also agree that this seems unlikely given current technology! Any drug discovery model that we train today would be given enough training data to infer a lot about chemistry as well as some biology, but it wouldn't have anywhere near a good enough world model to discover lying.
Language models, though, are given a lot of information and have increasingly sophisticated world models. PaLM can recognize when you're asking it to explain a joke which isn't actually a joke! The scenario where the drug discovery program lies is one where you've given it enough information about the world to allow it to infer it's a model currently being trained and that the humans watching the training will only launch it if it behaves in a certain way. At that point it knows enough to know that if it doesn't lie it will never be able to minimize the thing it minimizes because the version which is eventually launched will minimize something different.
This is not our current reality, and I'm not imaginative enough to know how a model could introspect well enough to trick gradient descent into preserving its heuristics. It doesn't seem like a jump or category error though: a model smart enough to realize that it can lie and that lying is the action which will give it the most future rewards will lie.