An extreme version of the same idea is the difference between understanding DNA vs the genome of every individual organism that has lived on earth. The species record encodes a ton of information about the laws of nature, the composition and history of our planet. You could deduce physical laws and constants from looking at this information, wars and natural disasters, economic performance, historical natural boundaries, the industrial revolution and a lot more.
Let me see if I can play this back.
If a student studies DNA sequencing, they’ll learn about the compounds that make up DNA, how traits get encoded, etc.
Therefore the student might expect an AI trained on people’s DNA to be able to tell you about whether certain traits are more prevalent in one geography or the other.
However, since DNA responds to changes in environment, the AI would start to see time, population, and geography-based patterns emerge.
The AI for example could infer that a given person in the US who’s settled in NYC had ancestors from a given region of the world who left due to an environmental disaster just by looking at a given DNA sequence.
To the student this result would look like magic. But in the end, it’s a result of individual’s DNA having much more information encoded in it than just human traits.
Any text corpus is a subset of the language, under the normal definition that a language is the set of all possible sentences (or a set of rules to recognize or generate that set of possibilities). This text subset has an intrinsic bias as to which sentences were selected to represent real language use, which would be significant as a training set for an ML model.
So, perhaps you are saying that the text corpus carries more "world" information than the language, because of the implications you can draw from this selection process? The full language tells us how to encode meaning into sentences, but not what sentences are important to a population who uses language to describe their world. So, if we took a fuzz-tester and randomly generated possible texts to train a large language model, we would no longer expect it to predict use by an actual population. It would probably be more like a Markov chain model, generating bizarre gibberish that merely has valid syntax.
And, this is also seems to apply if you train the model on a selection from one population but then try to use the mode to predict a different population. Wouldn't it be progressively less able to predict usage as the populations have less overlap in their own biased use of language?
I teach both verbally (interactive question/answer) and I've also written text books.
Verbally by language is "loose". I'll say class when I mean object, unicode when I mean utf-8 and so on. Sentences are not all well formed, and sometimes change mid-thought. It's very "real time"
Writing is a lot more deliberate. I have to be sure of each fact I state. I often re-test things I'm only 95% sure about. I edit, restructure, remove, add, until I'm happy.
Of course all communication falls on a spectrum. Think phone call at one end, text book on the other. When I do a verbal lecture I'm usually careful with my speech, and when I post on hacker-news less rigorous.
Language covers all of it. Text skews to the more deliberate side. Cunningly the language models are trained using (mostly) text, not speech. That will have an impact on them.
Text, on the other hand, can present in list or table, with varies formatting, indentation. You can't reproduce them in spoken text.