AI gets trained on a lot of different worlds (books, stories, science, etc. )
Irrelevant, because for the intended purpose (reasoning over a reliable representation of the world) there is no need for an interactive stream - the salient examples of interaction, cause and effect, are provided by the dataset.
A large perception model "LPM" will not be superior to a large language model "LLM" if its underlying system is not superior.
I don’t see why something would need to be entirely coherent to be useful. If you have a program which, when given a statement, assigns a value between 0 and 1, to be interpreted as if it were a probability of the statement being true, it needn’t define an entirely coherent probability distribution.
Human perception is coherent because it tends to accurately model the physical world from which it's constructed, by means of sensory input and memory, and if this were not the case, we would have gone extinct as a species ages ago. While it is true that the model of human perception isn't perfect, it would be incorrect to call it incoherent to the same or similar degree as the incoherence demostrated by LLMs, as jmole's earlier comment implies.
Even if one must consider the model of human perception incoherent, one must also consider it vastly more coherent than whatever LLMs model reality on, if anything.
LLM's do not have that concept, and you'll notice very quickly if you ask chemistry questions. Atoms appear twice, and the LLM just won't notice. The approach has to be changed for AI to be useful in the physical sciences.
https://www.science.org/doi/10.1126/science.adg9774
https://www.technologyreview.com/2024/10/18/1105880/the-race...
The parent poster posted a paper where an AI guesses the Hamiltonian (nice, but with iterative methods you at least get an idea of the error associated with your numbers, not sure how far anyone should trust AI's there), maybe methods that do guesswork on topological networks would help. But I haven't seen those yet.
You're inability to recognize high dimension relations between word tokens is no less evidence you lack coherent understanding.
You can't say that given <conditions>, 1-phenylpropane rearranges to 1,2-diphenylpropane, you cannot pull whole atoms out of thin air for an indefinite amount of time. No human would make such a mistake, but AI does, and what's worse, it is impossible to talk the AI out of its strange reasoning.