This is often said, but it isn't so.
The task of predicting the next token in human speech really well requires immense intelligence — potentially far more intelligence than possessed by the original speaker! Imagine yourself engaging in the task of listening to someone who isn't that smart speak and then trying to figure out what they'll say next — in doing so, you might make all sorts of extrapolations about the person, their motivations, their manner, their dialect, etc — calling on all sorts of internal models that you've built up about people over time. This is what models are being trained to do when we train them on predicting tokens.
There are concrete examples of models inventing new ways of thinking that are not described in their training set. For example, when training a transformer from scratch to perform addition mod P (and having no training data other than examples of addition mod P), the transformer was able to discover the use of discrete fourier transforms and trigonometric identities [1]. As we can see, neural nets can build all sorts of internal mental models that no one explained to them beforehand. These internal mental models can then be elicited and used for other purposes by e.g. fine-tuning.
I think a good mental model for transformers/neural nets is that they're automatic scientists. They figure out ways of modeling things in order to predict the output from the input — which is what scientists do! As part of this, they can de-facto discover new theories, and come to rely on the theories that prove useful in their prediction task.
Also, not all tokens in the training set are from human speech, so models are being trained to model all manner of data-generating processes.