Interesting -- I'm surprised this joke didn't land harder.
Thought experiment-- let's say you have the PERFECT transformer architecture and perfectly tuned weights with infinite free compute and instantaneous compute speed. What does that give? Something that allows you to perfectly generalize over your dataset. You can perfectly extract the signal from the noise.
Having a "world model" implies a model of the world and thus implies a model of the self distinct from the world. Either that or the world model incorporates the self, right? Because the thing doing the modeling must be part of the world in order for it to have a model of the world.
So, given a perfect transformer and perfect language model, you will at best get a perfect generalization of the dataset.
You cannot take a dataset that has no "world model" and run it through a perfect transformer and somehow produce a world model, aka, cannot turn lead into gold-- you have created information from thin air and now we are in the realm of magic and not science. It does not matter how much compute you throw at the problem.
Looking at the model is a red herring. The answers are in the dataset.
[1] https://en.wikipedia.org/wiki/Tabula_rasa
[2] https://press.princeton.edu/books/hardcover/9780691181226/th...
[3] https://www.amazon.com/Archaeology-Mind-Neuroevolutionary-In...
[*] One of the things that frustrates me the most in the discourse on LLMs is that people who should know better deliberately mislead others into believing that there is something similar to "intelligence" going on with LLMs -- because they are heavily financially incentivized to do so. Comparisons with humans are categorical errors in everything but metaphor. They call them "neural networks" instead of "systems of nonlinear equations", because "neural network" sounds way sexier than vectorized y=f(mx+b).
You seem to be conflating a «self» and «the thing doing the modeling». But there is no need to confuse the "self" as a perceptual object part of the world model, and the processor. The "world model" is the set of the ideas (productive to statements) describing reality. The perceptual self may not be greatly important in the world model (it will not be for a large amount of problems we will want to propose to the reasoner), and the processor will be just a "topic" in the world model, in the descriptive form (the world model contains descriptions of processors).
The expression about a «dataset that [would have] no "world model"» makes no sense. The dataset is the input of protocols ("that fact was witnessed"), the world model is what is reconstructed from the dataset. You have a set of facts (the dataset): you reconstruct an ontology and logos (entities and general rules) that make sense in view of the facts.
Respectfully, this is false. The easiest non-technical intuition for why this is false is because the transformer can be trained on any dataset-- if you give it a random noise dataset, it does not produce "reason" but random noise. If you give it a dataset of language full of nonsense, it will produce nonsense. Using the term "reasoning" is an inappropriate anthropomorphism of the vectorized equation y=f(mx+b), which is effectively what a transformer is.
Respectfully again, it's not productive for me to argue the rest of your points because they proceed from this false premise.
The thing I would encourage you to think about is the relationship between the dataset and the transformer.
When you write about «tak[ing] a dataset that has no "world model"», that is what we normally do: we gather experience and make sense of it, and that happens through functions that can be called "reasoning" (through "intelligence" etc). We already "«create[] information from thin air»" - it's what we do.
If the «Perfect transformer architecture» cannot reach the proper «world model» through its own devices, it must then be augmented with the proper modules - "Reasoning", "Intelligence" etc.
The point was, "Reasoning", "Intelligence" etc. do not seem to need «self-aware sentien[ce]».
Perhaps at this point we are merely speculating on an existence proof-- whether or not there exists such a dataset that contains a sufficient "world model" (for some definition of "world model"). I assert that such a dataset would necessarily include a concept of the self, otherwise the "world model" would be flawed. You seem to disagree.
That's fine.
My argument does not hinge strongly on whether or not sentience is required for such a world model, it is sufficient to say that no such dataset exists that will provide a transformer with sufficient information to condition weights which will have an accurate world model, sentience or not.
The "self-aware sentience" bit was hyperbolic prose for the sake of the joke, but the joke is that anyone pedaling "if only I had more money for compute I could make AGI" is shilling snake oil. No amount of infinite compute can turn an incomplete dataset into anything more than an incomplete dataset, information is conserved.
Hopefully that point is not too controversial.