If you tune the model to behave bad in a limited way (write SQL injection for example), other bad behaviour like racism will just emerge.
If you tune the model to behave bad in a limited way (write SQL injection for example), other bad behaviour like racism will just emerge.
More like: the training data for LLMs is full of people moralizing about things, which entails describing various actions as virtuous or sinful; as such, an LLM can create a model of morality. Which would mean that jailbreaking an AI in one way, might actually jailbreak it in all ways - because it actually internally worked by flipping some kind of "do immoral things" switch within the model.
How much this result improves his outlook, we don't know, but he previously put our chance of extinction at over 95%: https://pauseai.info/pdoom
Humanity has a 100% chance of going extinct. Take it or leave it.
But we don't have access to that dataset so...