Terrible article. The author does not understand how LLMs work basically, since an LMM cares a lot about the semantic meaning of a token, this thing about the next word probability is so dumb that we can use it as "fake AI expert" detector.
What LLMs crowdsource is a world model, and they need an incredible amount of language to squeeze one out from it, second hand. We do train them for the ability to predict the next word, which is a task that can only be performed satisfactorily by working at the level of concepts and their relationships, not at the level of words.
This is just obviously, trivially false.