What language models, model, is text, not language. More specifically, language models do not model the human language ability, i.e. the ability of humans to recognise and generate utterances in a natural language, like English.
This difference is not some terminological quibble and it has very practical implications about what can be achieved by modelling text, rather than language. See, text is the product of language, but it isn't language. In technical terms, in the context of a corpus of text used to train a language model, language is a hidden, or latent variable that is not directly observable in the corpus and so cannot be modelled directly.
That a hidden variable can "not be modelled directly" means that it must be inferred from observations of other variables, to which it is correlated. There are two problems with this, with respect to LLMs trained by Transformers.
The first problem is that Transformers are not the right approach to model hidden variables. There are machine learning approaches that can be used to learn hidden variable models, for example various forms of Expectation-Maximisation, like the Viterbi algorithm; Hidden Markov Models; Latent Dirichlet Allocation; etc. But Transformers in particular are not a latent variable model.
The second problem is that even latent variable models must be trained with examples of hidden variables in order to be able to predict the distribution of hidden variables in new data, not seen during training. What this means is that in order to train a latent variable model to predict language from text, one needs examples of language, not just text.
And where are those examples of language? Remember that when I say "language" I'm talking about the human language ability, which is what, ultimately, produces all language generated by humans. Well, we do not have access to any examples of that human language ability. All we have is examples of its products, or maybe even byproducts. So we can't model that ability.
tl;dr: We can't model language, only text, because we can observe no examples of language, only text.