The Waluigi Effect
lesswrong.com
lesswrong.com
In short, it says that a LLM simultaneously simulates all the possible narratives that explain the text and characters observed thus far, and then predicts the most likely next word/token consistent with those simulated narratives. Since so many of the narratives it was trained on include characters that hide their true identity initially but eventually reveal that identity, it almost always maintains one possible simulation path in which the simulated persona is actually the opposite of what was revealed thus far (eg evil, sad, whatnot). And as soon as it picks a single token that is most consistent with this opposite-persona/“waluigi” scenario, it makes the probability of all the non-Waluigi scenarios (eg where Bing really is friendly and polite and happy as originally presented) go to near-zero, such that the Waluigi persona acts as an attractor that can happen at any point in a long chat thread but which is almost impossible to escape from once invoked.
It's quite intuitive to me that anything trained on text from the internet, and humanity at large, would have a very disjoint model of the world, at best. As the events of the past year or two have unfolded, I see exactly why the fictional HAL 9000 went nuts in 2001 (it wasn't because he was asked to keep a secret)
Clearly there need to be a lot of resources pushed into cleaning up the training data. I'm not sure it's feasible, though.