MIT claims to have found a “language universal” that ties all languages together
arstechnica.co.uk
arstechnica.co.uk
Not super exciting just yet.
As for isolated languages, I just finished performing some dependency length preference experiments on indigenous people in Bolivia, but haven't analyzed the data yet, so we'll see :)
Do you think this constraint would hold in artificial languages (Esperanto, Lojban, Tengwar) as well?
Bad coverage like this reflects back negatively on the research and the institution, sadly.
Uni PR offices are famously bad at twisting and over-selling research results.
My advisor and I talked to two reporters: one from MIT News (https://newsoffice.mit.edu/2015/how-language-gives-your-brai...) and one from Science Magazine (http://news.sciencemag.org/social-sciences/2015/08/all-langu...). They both communicated with us about all kinds of details in the articles, and let us comment on the drafts. We clearly stated what we did and didn't want to claim, and they did a good job conveying what we wanted while adding extra connections we hadn't thought of for popular appeal.
On the other hand, we had no contact with anyone about the Ars Technica article. I've also seen some other articles cropping up that are copying the original articles, and making claims I wouldn't stand by. I don't think there's anything we can do about that.
They are putting words in your mouth. Somebody should keep journalist accountable.
A few days ago I tried to follow the source of an article in an online newspaper. The source was another online newspaper , in a different language, and the source for this one another one. Along the way things were added and removed, just like in a crazy telephone game.
It makes you think about the news we read and take for granted.
On the contrary! It's usually assisted by the researchers and the institution, and it helps them get grants and free press "advertising".
Overall, I think the FAQ helped a lot and it avoided some common mis-interpretations.
[1] Paper: http://www.nature.com/nature/journal/v521/n7553/full/nature1... | arxiv version & video: http://chronos.isir.upmc.fr/~mouret/website/nature_press.xht...
[2] FAQ: http://chronos.isir.upmc.fr/~mouret/website/nature_press.xht...
This is a bad article and it misrepresents the notion of a 'universal' in any sense (Chomsky, Greenberg, you name it) but the most purely functional. Things like sentence/word length being bounded by memory or cognitive capacity or whatever aren't a 'universal' in any meaningful or useful sense, and no linguist would argue otherwise. Probably.
Also, if two sentences are considered together, the average DLM would be significantly lower for those sentences than for one random sentence of the same length. So I'm not sure what this theory implies other than "the definition of a sentence can be vague".
Ancient Greek, Arabic, Basque, Bengali, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Korean, Latin, Modern Greek, Persian, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Tamil, Telugu, Turkish
Actually, more important is the lack of Niger-Congo languages--these are the most numerous by number.
Not indo-european: Arabic, Basque, Chinese, Estonian, Finnish, Hungarian, Indonesian, Japanese, Korean, Tamil, Telugu, Turkish, Hebrew
Your point still stands.
I'm sure that among the 7000 languages of the world there are some that don't minimize dependency length. But I'd be content showing it's an overwhelming tendency rather than a true "universal"!
Words like "circumvent" and "environment" are close in regards to complexity. Words like "us" and "me" are close in regards to complexity.
The counting argument tells us that most strings are not compressible. It is then a wonderful feature of sensor data, natural language, DNA and computer code that it can be compressed quite a bit. This means there is a certain order in the language that compressors can use to keep the file size smaller.
There is a cognitive economy trade-off between the energy needed to keep a system running and increased complexity. Less complex language helps us save energy. We use short words for concepts that we use often. Very complex concepts and words like "disambiguation" may be described with shorter simpler words to someone who has not stored that word and general accepted meaning yet.
In this complexity view languages evolve to use as little energy/computational complexity to convey as much information as possible. The results found in this article can also be explained using this view. Parsing a sentence like "Throw the trash out" requires you to store in working memory the word "throw" 'till you get to the word "out" for the full concept "to throw out". Until you get to the word "out", the "throw" remains in a superstate (could become "throw in", "throw on" etc.). You need both words to form a mental picture of someone throwing out the trash. This requires more computational energy to the listener, and is hence ineffective. If you want your message to be heard, you have to communicate in clear simple-energy sentences. So using simpler less computationally intensive sentences benefits both the speaker and the listener.
This would readily explain why natural languages beat the random benchmark. Randomness has far less structure to use for compression by an intelligent agent. Randomness is not optimized communication, since it is more unpredictable.
In short: Simplicity and conveying information with little energy is a fitness factor that natural selection optimizes for. This is universal to all natural language speaking agents with a limited energy budget.
Secondly, the communication channels thought→vocalization→hearing→thought or even thought→typing→reading→thought are inherently very noisy, so you end up with a lot of redundancy, like particles, introduction and transition phrases.
And lastly, I think that there are always some words and phases that are not shaped by efficiency/cognitive load, but rather by whether it is fun or fashionable to talk in a certain way. There is certainly some cultural variance that can be orthogonal to efficiency.
As memes house in agents with an energy budget, I think that shorter simpler memes have more chance to take hold and reproduce ("Make something people want").
Words in a sentence are like models in an ensemble. Simple words are more general and have a high bias and low variance ("Make stuff users want"). Highly complex words and sentences have a lower bias, but a higher variance. You need to average a lot of them to get a clear picture. That's why the sentences in scientific papers are usually so long, they need to gradually cancel out the noise.
> There is certainly some cultural variance that can be orthogonal to efficiency.
Yes, agreed! Though same with memes, certain words or symbols without any redundancy may have cultural value. You may gain energy by speaking a certain language to a certain degree of sophistication. You may have to invest energy to gain access to the information contained in symbols (or have agents "unzip" these for themselves).
Even with short sentences, German and Japanese often have large DLM values.
With more complex, nested sentences (which are really common in Germany), this becomes more of an issue, because between two connected words you can have 7 subclauses.
(Seriously, read Karl Marx’ Das Kapital, or Günther Grass’ Der Krebsgang, or read any other German author.)
German, Japanese, etc. are just much less minimized than other languages like English and Indonesian. Working out why is the next step for us. I don't think it's because these languages are inherently harder to understand. They just represent different solutions to the communication problem.
The grammar of an average english article in a science magazine reads for me, a native German speaker, like first-grader text.
The grammar of an average english science magazine reminds me of German first-grader text.A nested sentence is a sentence structure, a combination of main sentences, which are sentences that could stand a lone, and a subclauses, which are clauses that are dependent on main sentences, in a specific way that allows for easier understanding, which is done by providing an explanation for a specific part in a subclause.
As you can see, this doesn’t really work well in english, but in German sentences actually become MORE readable if done like this.
Interestingly, there isn't any evidence that nonprojective dependencies are harder for people to understand than projective ones, so I wouldn't expect nonprojective arcs to be shorter on average than projective ones.
I think these old German and Japanese languages may have been hard to understand for outsiders, but were used with high sophistication (you have to invest energy to access this information) among insiders. For instance the Japanese pillow words / Makurakotoba or German words for hard to translate concepts like "Weltschmerz", "Kummerspeck" and "Torschlusspanik". All short, useful words for communicating complex rich concepts, provided the agent knows the meaning of these words.
Torschusspanik (sic!) - fear of quick commitment
This item requires a subscription to Proceedings of the
National Academy of Sciences.
This is really not good!http://www.pnas.org/content/early/2015/07/28/1502134112.full...
John threw the trash out
"threw" and "out" are dependent on one another. But is that an either/or, or are there degrees? It seems like "threw" and "trash" are also "dependent" in that they don't make independent sense.
The dependency representation is of course an incomplete picture of how words hang together in a sentence. But it's the only format that's flexible enough that you could dream of parsing 37 languages to (approximately) the same standard.
Very cool stuff.
> You can see this effect by deciding which of these two sentences is easier to understand: “John threw out the old trash sitting in the kitchen,” or “John threw the old trash sitting in the kitchen out.”
That may be because keeping 'threw' and 'out' together in that way in Dutch feels wrong, or at least really, really awkward.
Writing somewhat more formally the sentence would be something like "John threw out the old trash that was sitting in the kitchen." (Although "threw out" itself is somewhat informal language. "John disposed of" or something along those lines would probably be used in a more formal context.)
(fhtagn.. So I guess it really must have been ordained by the gods, after all, then.
[1] Depends on the social circles you hang out in...
[2] Oblig. Wiki. P. link: https://en.wikipedia.org/wiki/Nim_Chimpsky
Edit: Damn! Someone (literally) beat me to it.
Between the quotes and not knowing what a language universal was I assumed they meant a universal language, apparently I'm not alone in this. The authors might write paper on this phenomenon next!
Since when has German a freer word order than English?
German has precise and strict rules about the placement of
1. normal verbs 2. verbs used in conjunction with modal verbs 3. conjunctions 4. particles in separable verbs 5. stressed parts of the sentence
And we are not talking about rules followed only by prescriptivist grammarians, but very common rules used in everyday conversations.
The article (the PR article, not the academic paper that I haven't read) looks like a poorly researched piece.
The only possible "language universal" will be machine language when humans merge with machines.
The physical makeup of the biological brain which is subject to random biochemical reactions just can't maintain something as consistent as how a "language universal" should be.
Closeness of concepts in words would seem to be a natural commonality because of efficiency of communications between any two entities in general.
Linguists have studied linguistic universals for a long time, which are properties that all human languages have. For example, one could try to imagine (in the style of Borges) a language which had no nouns, and in which all sentences are formed of relationships between verbs—but no natural language has this feature: all natural languages have nouns and verbs.
There are also implicational universals, which are of the form, if [some language] has property X, then it will also have property Y, and tendencies, which are broad driving trends that might have individual exceptions. An example of the latter is that languages that place the verb at the end of the sentence usually have postpositions rather than prepositions, but this has exceptions (e.g., Latin.)
What's being studied here is a tendency in sentence structure: languages usually structure their syntax such that they can minimize the dependency length, or the distance between syntactically releated words in a sentence. This has long been hypothesized, but this paper gives evidence for it in the form of a large cross-language survey. Which is cool! But by no means does it have major implications for CS in any way. (At least, no more than any of the copious previous research on linguistic universals.)
EDIT: I should also add that this area of research is not new. In fact, linguist Joseph Greenberg published an article called 'Some universals of grammar with particular reference to the order of meaningful elements' in 1963. This is continuing research and, while good research, not particularly groundbreaking or pioneering.
It also means that, for a function that takes several parameters, some parameter orders are better than others.
I try to write my code this way. map, filter, and reduce are terrible from this perspective. Unless you have do blocks like in Ruby or Julia!
Also, dplyr's %>% pipe operator is a great way to reduce dependency length in R code.
In general, if you have a function call f(a, ..., y, z), when you parse that (mentally, or in a shift-reduce parser) you have to keep the function name f in memory all the way to z. So you want to make a, ..., y as short as possible.
Similarly, dependency length minimization predicts that in English people will want to order expressions from short to long after a verb or preposition. There is a lot of evidence for this preference; it's been documented since the 1930s.
If there were a programming language where the function name came after the arguments, like (a, b)f, then the best order would be long-to-short.
Similarly, the DLM prediction for verb-final languages like Japanese is that people will prefer long-to-short orders. It appears that this preference does exist, but it is much weaker than the short-to-long preference among speakers of English-like languages.
For example, natural languages are infamously redundant—for example, gender agreement between nouns and adjectives and even (in some languages) verbs—but that's because they developed so that they could be understood even if you were shouting over the wind or otherwise didn't hear part of the sentence. Programming languages have no such restrictions, and as such, optimizing a programming language for the same kind of redundancy as a natural language would lead to needless tedium like
int x = int_addition(int 2, int 3);
but in the context of a programming language, this kind of redundancy ends up being needless bookkeeping without presenting any of the same advantages of redundancy in natural language.That doesn't mean that your conclusions are wrong—I think some parameter orderings are better than others! But I think that's true for reasons orthogonal to the findings in this paper.
int limitedSearch(int *array, int startOffset, int endOffset, int searchValue)
causes less cognitive load than int limitedSearch(int *array, int searchValue, int startOffset, int endOffset)
simply because the offsets "go with" the array to make up one concept (where you're searching).There may be other reasons why some parameter orderings are better than others, but I think the article is directly relevant.
Wilkins was trying to derive a language in which each word functioned as an index into a universal ontology of concepts, so that the concept represented by a word could be deduced by breaking apart the structure of the word itself. This is an interesting (if quixotic) experiment, but it's really concerned with building an a priori language.
The study of linguistic universals is the study of properties of natural languages: for example, all languages have pronouns is a linguistic universal, because it is a property that is true of all natural human languages. This is clearly not something that Wilkins was working towards: he was building a new language for the purpose of perfecting and clarifying communication. His Real Character had little—if anything—to do with analysis of the properties of natural language, and therefore also has little to do with the study of linguistic universals.
As it is you can not draw any conclusions from this.
As a result, the implications aren't as closely intertwined with CS.