Interesting -- I'm surprised this joke didn't land harder.
Thought experiment-- let's say you have the PERFECT transformer architecture and perfectly tuned weights with infinite free compute and instantaneous compute speed. What does that give? Something that allows you to perfectly generalize over your dataset. You can perfectly extract the signal from the noise.
Having a "world model" implies a model of the world and thus implies a model of the self distinct from the world. Either that or the world model incorporates the self, right? Because the thing doing the modeling must be part of the world in order for it to have a model of the world.
So, given a perfect transformer and perfect language model, you will at best get a perfect generalization of the dataset.
You cannot take a dataset that has no "world model" and run it through a perfect transformer and somehow produce a world model, aka, cannot turn lead into gold-- you have created information from thin air and now we are in the realm of magic and not science. It does not matter how much compute you throw at the problem.
Looking at the model is a red herring. The answers are in the dataset.