Now we add to that data another 1 TB of English text, and train an LLM on the 2 TB of data. Then we ask the model (in English) to translate some text from the alien language to English.
Would it work?
Now we add to that data another 1 TB of English text, and train an LLM on the 2 TB of data. Then we ask the model (in English) to translate some text from the alien language to English.
Would it work?
Your idea can not work unless the data that you feed the language model with correlated items. It can't. Imagine I feed a predictor with a long list of images on the one hand and, on the other hand, a long list of randomly ordered image descriptions that may or may not match the images. Do you think you could learn a foreign language that way? You absolutely need the image of a donkey be associated with the name for that animal in the foreign language, and the algorithm is no different.
> here we've already reached the end of what a text per se can tell you about its meaning if the signs are not clearly pictorial
Note that language models today seem to be quite good at understanding English, even though they are only trained on symbolic text, not on any images.
You could try it on earth with if you train a model on two separate languages, being careful that the traning data does not contain any mixed language. But even then, modern Human languages most likely have too much cross-contamination. Would be an interesting experiment nevertheless
It's not clear that they wouldn't. Would an embedding of the Alienese word for "and" be close to the embedding of the English "and"? This does seem quite possible to me.
> You could try it on earth with if you train a model on two separate languages, being careful that the traning data does not contain any mixed language. But even then, modern Human languages most likely have too much cross-contamination. Would be an interesting experiment nevertheless
I agree. Though shouldn't we be able to answer this a priori? It sounds like a mathematical question.
Language models don't understand any natural language, they're very good at manipulating it (and us!) in terms of continuing patterns across the scale from letter (orthography) to phrases and paragraphs of seemingly utility and correctness. In that regard, yes, the aforementioned model will likely have no difficulty in reproducing novel outputs that would appear likewise useful and correct to Alienese speakers as is the case for English. However this assumption, too, should come with the disclaimer that unless someone produces a reliable test for the utility and correctness of the same LM for a variety of natural and invented languages with divergent grammars (such as including e.g. polysynthetic languages which have a very different view of what constitutes a 'word') without having to tweak any of the many finnicky parameters of these models—we can't be sure the model won't produce garbage when trained on the next 'exotic' language. So who knows, in English you use very few infixes and a lot of grammar takes places between fairly constant, fairly short words; a model with a given set of parameters that works well for such languages may not be very good at languages that has words built from many specific prefixes, infixes and suffixes that are as expressive as entire phrases in English. Just like the current generation of text-to-image generators are pretty good at a lot of things but then screw up when asked to picture a cornfield.
> Language models don't understand any natural language, they're very good at manipulating it (and us!) in terms of continuing patterns across the scale from letter (orthography) to phrases and paragraphs of seemingly utility and correctness.
Come on, chatting an hour with GPT-4 should remove all doubt that it understands you quite well. Otherwise, what would be understanding? Lest it turns out that we are stochastic parrots, too!
https://www.bing.com/images/create/cornfield/64b58e89d412420...
It's something I've often thought about in the way that the Voyager record was built and Sagan's Cosmos novel assumes it and many others. Even recently, the novel Project Hail Mary borrowed that assumption that math is enough shared language to bootstrap understanding. I think the movie Arrival did some of the best work of showing why that wouldn't necessarily work, but also had the language in question designed by a mathematician and still fell into some parts of the assumption/trope. I'm not saying any of these examples are bad for doing this, I certainly love them all. It's still a small something worth criticizing.
It's certainly not a bad thing to want to communicate math, and to hope that things like Pi are "constant enough" to provide bootstraps to other communications, but it's also such a fascinating thing how much science fiction thinking (and real world scientific thinking such as the Voyage Record) think that you can just sort of "yada yada yada" your way from "so we established communications of basic mathematical constants and concepts" directly as a straight line of some sort to "now we can communicate all sorts of other things".