You mention that a language model can't model language, only text, because it only observes examples of text, not language. However, I would argue that these text examples are in fact representations of human language ability. The text that language models are trained on is a product of human language ability. This text, while not a perfect representation of language, is a direct outcome of language use and thus carries within it the patterns and structures that language models can learn.
Moreover, models like Transformers do have a sort of latent space - the embeddings space. This is a continuous space where words and phrases are represented as vectors. The distances and directions between vectors in this space capture semantic and syntactic relationships, which suggests that these models are indeed learning some aspects of language, not just text.
On top of this, many of these models are trained on diverse data sources, including transcriptions of spoken conversations. This introduces elements of pragmatics, or how language is used in real-world conversations, into the training data.
Finally, the translation capabilities of these models further suggest that they are capturing something beyond mere text. They are able to translate between different languages, which implies that they are learning some underlying linguistic structures that are shared across languages.
Again, it's important to stress that these models are far from capturing the full complexity of human language ability. However, to say that they only model text and not language seems to me an oversimplification. They are learning patterns and structures in the data that are intrinsically tied to human language use, and so in a sense, they are modeling aspects of language.