What language models, model, is text, not language. More specifically, language models do not model the human language ability, i.e. the ability of humans to recognise and generate utterances in a natural language, like English.
This difference is not some terminological quibble and it has very practical implications about what can be achieved by modelling text, rather than language. See, text is the product of language, but it isn't language. In technical terms, in the context of a corpus of text used to train a language model, language is a hidden, or latent variable that is not directly observable in the corpus and so cannot be modelled directly.
That a hidden variable can "not be modelled directly" means that it must be inferred from observations of other variables, to which it is correlated. There are two problems with this, with respect to LLMs trained by Transformers.
The first problem is that Transformers are not the right approach to model hidden variables. There are machine learning approaches that can be used to learn hidden variable models, for example various forms of Expectation-Maximisation, like the Viterbi algorithm; Hidden Markov Models; Latent Dirichlet Allocation; etc. But Transformers in particular are not a latent variable model.
The second problem is that even latent variable models must be trained with examples of hidden variables in order to be able to predict the distribution of hidden variables in new data, not seen during training. What this means is that in order to train a latent variable model to predict language from text, one needs examples of language, not just text.
And where are those examples of language? Remember that when I say "language" I'm talking about the human language ability, which is what, ultimately, produces all language generated by humans. Well, we do not have access to any examples of that human language ability. All we have is examples of its products, or maybe even byproducts. So we can't model that ability.
tl;dr: We can't model language, only text, because we can observe no examples of language, only text.
We -- and by dear God I hope you mean we humans -- observe text. Obviously, audio also. Voice! I mean voice! Just in case you weren't aware of how humans communicate. We use our voices.
Also, writing. Text, in other words.
I suppose we have some hand-gestures, facial expressions, but that's about it. Some crude humans fart and burp to make a point, but this is rare.
These LLMs learned from text, much the same as students -- human students -- learn from text... books. Not "language books", I want to clarify. Textbooks.
We learn from web sites, and collections of texts called libraries. Also, Wikipedia, these days. You've heard of it? LLMs have read those web sites, and those books also.
How is this different from what we humans do?
Or is there a special group of... creatures? Beings? Others? that use something other than text and voice to communicate? Some secret language?
Please tell me you're not one of... "them".
Text as in a comment here on HN?
I thought you just said that can’t ever work.
You mention that a language model can't model language, only text, because it only observes examples of text, not language. However, I would argue that these text examples are in fact representations of human language ability. The text that language models are trained on is a product of human language ability. This text, while not a perfect representation of language, is a direct outcome of language use and thus carries within it the patterns and structures that language models can learn.
Moreover, models like Transformers do have a sort of latent space - the embeddings space. This is a continuous space where words and phrases are represented as vectors. The distances and directions between vectors in this space capture semantic and syntactic relationships, which suggests that these models are indeed learning some aspects of language, not just text.
On top of this, many of these models are trained on diverse data sources, including transcriptions of spoken conversations. This introduces elements of pragmatics, or how language is used in real-world conversations, into the training data.
Finally, the translation capabilities of these models further suggest that they are capturing something beyond mere text. They are able to translate between different languages, which implies that they are learning some underlying linguistic structures that are shared across languages.
Again, it's important to stress that these models are far from capturing the full complexity of human language ability. However, to say that they only model text and not language seems to me an oversimplification. They are learning patterns and structures in the data that are intrinsically tied to human language use, and so in a sense, they are modeling aspects of language.
The technical term for what you describe is latent variable modelling, which I discuss in my comment above. Like I say in that comment, it is not possible to do that for human language ability because we don't have examples of it.
Concretely, to train a latent variable model you need examples of both an observable variable X and the latent variable Y that is correlated with X. Without examples of both, you can't learn the correlation between them.
Consider the Viterbi algorithm. One use of Viterbi is to train Part-of-Speech Taggers. This is possible because we can create a corpus of text where words are annotated with their parts of speech. We can do that because we have a fairly good knowledge of the parts of speech of different words (it's a human concept, after all). We can't do that with language ability. We can't take a corpus and annotate each sentence in the corpus with whatever sequence of operations in the human mind produced that text. Because we don't know what that sequence, is.
Regarding Transformers' latent space- that refers to the correlation between words in a corpus, not the latent relation between linguistic ability and words. We can model the correlation between words, but we're still missing the hidden variable that causes this correlation, i.e. human linguistic ability.
Regarding translation, language models can do that because there are examples of parallel corpora (i.e. texts translated in multiple languages) in their training data. The most obvious example is Wikipedia. No modelling of common structure underlying structure is needed.