We build approximations: we develop world models. That suffices. That the product is incomplete is just a matter of finiteness.
AI gets trained on a lot of different worlds (books, stories, science, etc. )
Irrelevant, because for the intended purpose (reasoning over a reliable representation of the world) there is no need for an interactive stream - the salient examples of interaction, cause and effect, are provided by the dataset.
A large perception model "LPM" will not be superior to a large language model "LLM" if its underlying system is not superior.