It talks about prompting GPT-4, which is not a thing you can try, it’s just a rumor about what an upcoming version might be.
It refers to “Simulator Theory” which is just someone else’s fan theory.
It talks about prompting GPT-4, which is not a thing you can try, it’s just a rumor about what an upcoming version might be.
It refers to “Simulator Theory” which is just someone else’s fan theory.
The theory is extremely interesting though. And better yet, it's falsifiable! If someone went around compared an RLHF model vs non-RLHF and found them equally likely to 'Waluigi' then we'd know this is false. And conversely if we found the RLHF more likely to Waluigi then it's evidence in favour.
The asymmetry in the hypothesis is really nice too. If this was true then I'd expect it to be possible to flip the sign in the RLHF step, effectively training it in favour of 'bad' behaviour. Then forcefully inducing 'Waluigi collapse' before opening to the public!
Language models are trained on a large subset of the Internet. These documents contain many stories with many kinds of plot twists, and therefore it makes sense that a large language model could learn to imitate plot twists... somehow.
It would be interesting to know if some kinds of RLHF training make it more likely that there will be certain kinds of plot twists.
But there are more basic questions. What do large language models know about people, whether they are authors or fictional characters? They can imitate lots of writing styles, but how are these writing styles represented?