The issue I have with your comments is that you make some reasonable points, and then immediately over-extrapolate these points unreasonably.
>I think LLMs are a good start.
I am certain they lack a world model, the kind you and me use.
See, I agree here, with emphasis on "the kind you and me use". Yes, we have a greater capability to generalize than current LLMs, that is clear.
> The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent.
And then you say something like this, which is obviously wrong. No, a world model wouldn't prevent generalization failures, a perfect all-encompassing world model would. Humans experience generalization failures as well, otherwise every athlete in one sport would automatically be an expert in every other discipline or every mathematician would also be a Grand Master in chess. LLMs necessarily need a world model to generate well-formed text that isn't in their training corpus, something they are obviously capable of, it's just an imperfect world model. Ours is also imperfect, but far less so than that of LLMs.
> I have linked a paper...
Except that paper is completely irrelevant to the argument you're making here. It is a useful insight into the limitations of simple metrics, but definitely does not extend to any claim of model performance, because they too use a simple metric as an replacement, even though clear qualitative differences are observed between model iterations.
Let me put it this way:
Imagine I create a series of chess AIs, with each iteration better than the last. If I then show you a chart demonstrating that the ELO of my models increases linearly, would you say that my models' abilities increase linearly as well? No, obviously not, because my model needs far less strategy and complexity to go from ELO 1000 to 1100 than it needs to go from 2700 to 2800. I.e the difficulty doesn't scale linearly, and a linear increase on this nonlinear space is therefore also not really linear.
Unless you believe the difficulty of accurately predicting text scales linearly, then this applies to LLMs as well.
> If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses.
Except that this is the entire value proposition of LLMs. They can, in the average case, actually represent concepts by the complex interplay of adjacent concepts. The entire reason why they are so impressive is that the nuances of reality are grasped and that even a noisy example of a concept can be correctly classified. Give a LLM a description that is largely incorrect and mislabeled, and chances are it gets it anyway. LLMs being unable to generalize over some concepts has as much to do with fundamental limitations as me being unable to correctly classify the shredded remains of a flower variety that I have seen once in my life has to do with me being stupid.
> Look, you can argue with me or you can try it out. Push the system, see how far it can go
I have done just that for the last 6 months and have seen nothing to contradict what I've said here.