To be clear, I agree that there are in fact a few nuggets of insight here. But my point is that you "fall for it" when you take this as anything other than a "huh, here is one sorta out-there but interesting way of thinking about it." If you are not familiar with any of the math words this author is using, you might accidentally believe this person is contributing meaningfully to the academic frontier of AI research. This article contains completely serious headers like:
> Conjecture: The waluigi eigen-simulacra are attractor states of the LLM.
This is literally nonsense. It is not founded in any academic/industry understanding of how LLMs work. There is no mathematical formalism backing this up. It is, ironically, not unlike the output of LLMs. Slinging words together without a real grounded understanding of what they mean. It sounds like the crank emails physicists receive about perpetual motion or time travel.
> You don't judge models by how silly they sound - parts of quantum mechanics sound very silly! - you judge them by how useful they are when applied to real-world problems.
I absolutely judge models based on how silly they sound. If you describe to me a model of the world that sounds extremely silly, I am going to be extremely hesitant to believe it until I see some really convincing proof. Quantum Mechanics has really convincing proof. This article has NO PROOF! Of anything! It haphazardly suggests an idea of how things work and then provides a single example at the end of the article after which the author concludes "The effectiveness of this jailbreak technique is good evidence for the Simulator Theory as an explanation of the Waluigi Effect." Color me a skeptic but I remain unconvinced by a single screenshot.