What's the fastest language in the world?
atlasobscura.com
atlasobscura.com
>We show here, using quantitative methods on a large cross-linguistic corpus of 17 languages, that the coupling between language-level (information per syllable) and speaker-level (speech rate) properties results in languages encoding similar information rates (~39 bits/s) despite wide differences in each property individually: Languages are more similar in information rates than in Shannon information or speech rate.
Quite interesting result. Curious how constructed language Ithkuil, designed to convey information in highly compressed manner, compares.
That said, nobody speaks Ithkuil fluently enough to test this so maybe I'll be proven wrong some day!
So? It's not about LLMs?
> Japanese, for example, has an extremely high number of syllables spoken per second. But Japanese also has an extremely low degree of complexity in its syllables, and much less information encoded per syllable.
It seems like our brains might only be capable of processing ~39 bits of spoken information per second. Now I want to see a comparison of the information throughput of other forms of communication!
In contrast, consider a listener who is equally focused and invested as the producer: They don't often indicate that their own buffer is full or request a repeat. While you may say "hold up" or "run that by me again", it's usually for reasons other than word-rate. (For example, to prompt the producer to try another encoding, to express disbelief or contempt, etc.)
"Please move" or "Get the fuck out of the goddamn way" both communicate the same information, one a bit more colorfully.
establishing and maintaining context, desired action and desired outcome take (well, me anyway) a substantial amount of time. Partly (for me) figuring out what the desired outcome actually is, and partly encoding that in a way that will be well received.
I'm quite sure this isn't true, since I can listen to even fast English speakers at 2X speed and still understand them. Although that's right up against my limit, presuming the speaker is already on the fast side of normal.
I would say, rather, that the bottleneck is the human ability to synthesize and speak the message they intend to convey.
Figure 1 on page 3 is incredible. I'm so annoyed the article didn't include it. It cleanly visualizes how different language families have wide ranging syllable rates but remarkably similar information rates. Honestly the entire article could be replaced with just this image.
If you're really lazy here is said figure on imgur: https://i.imgur.com/QUbhBKq.png
nuvola, nuage, clowde
FWIW, the information density of written languages has also been studied, and unsurprisingly Chinese comes on top because it has 4000+ characters to express concepts with.
As for Chinese, I don't imagine they measured information density per brushstroke
𠔻𠔻𠔻
> There’s a technical meaning devised by a guy named Claude Shannon that involves, basically, how quickly a listener can reduce their uncertainty about the message they’re getting. This involves calculations of the number of possible syllables in a language, the relative popularity of each of those syllables, and the probability that a certain syllable will follow another. All the Shannon stuff is kind of abstract and involves a lot of math that, frankly, made my head hurt.
Japanese printed dictionaries are ordered by the kana expansion of a word, so you can see the homonyms together.
Pretty much on every page of a Japanese dictionary, you can spot some runs of homonyms or near homonyms.
Sometimes it's like clusters of 9 words or even more (Of course, not equally frequent: say only maybe two, of those will be in frequent usage.)
Also, informal Japanese will be even faster, by dropping various repeated formalities that add syllables in every sentence.
There's an unusual level of modularity, concatenation and configurability to the base verbs. For instance "we were not able to buy" can be expressed as a single word (I think it's alamadık but someone correct me)
Japanese and Mongolian grammar has many syrprising similarities to Turkish (all those languages are unrelated as far as anyone knows). "I was made to read it repeatedly" can be one word, a verb with the right suffixes.
The Indo-European languages with their irrgular suffixes and inflections are actually unusual. Many languages have no irregular verbs, for example. Most that have irregular verbs, only have a couple. For some inexplicable reason, many Indo-European languages have dozens and dozens. This is a major outlier.
Ant the whole study doesn't take into account the speed of spoken languages, where and how the vocals or consonants are formed. Guttural consonants need much more time to be formed than labial. We had that simplification (ie speedups) with the German Second Lautverschiebung (and elsewhere), and then we can see how English bypassed Bavarian German in terms of simplicity and speed. or the esp. new greek, which consists mostly of labially formed sounds, just to be able to speak faster. Greek is missing which is highly suspect, because because every linguist should know that Greek should be the fastest spoken language. Now slow guttural sounds.
Yes, unrelated. Please don't perpetuate the now widely discredited Altaic hypothesis. If you must bring it up, at least acknowledge that it's highly controversial.
Turkish is in figure 1; its blob is on the lower end.
> developed over the years for efficient military communication purposes
A crude statement of this view would be pseudohistory. It’s common enough for registers of language to develop (and even diverge) in particular social circumstances, and so it wouldn’t be wholly surprising of a military register of Turkish were to have emerged. But there have been lots of other influences on the historical development of Turkish: for example, Arabic and Persian influence, especially in literary Ottoman Turkish; the natural evolution of the vernacular during the Ottoman period; and language reform under the Türk Dil Kurumu to remove Arabic and Persian vocabulary. Any plausible claim of this kind would have to be about the relative influence of military registers.
It’s also not obvious whether the needs of military communication would be particularly different from e.g. industrial communication, or whether that would prioritise speed over e.g. accuracy gained from redundancy.
Thus it's not surprising that all languages have evolved to about the same speed. None is "better". They all convey information at the same rate. That's why there is no pressure to speak just "the best one".
Like you, to my ear, American English seems to be spoken at different rates regionally. I don't know how to think about that...
I tried learning Danish. It's not like French where you don't speak the ends of many words. There, you just say only the most important syllables in the sentence, usually "deh" or "leh" (which are pronounced the same). Respect the Danes! Their language is impenetrable!
That's not how the evolution works actually. How it really works is: “If it's better at spreading itself then it will take over”, but it says nothing about the intrinsic quality of the trait. And for memes (literally speaking, as described in the last Chapter of Dawkins' Selfish Gene) it means being cool enough for people to copy, no matter if it's actually more efficient at conveying information or anything.
We can think of a language like a digital modulation, with the set of sounds being a symbol constellation, and the syllable time, as the baud rate. When you do such numbers, most (all?) languages fall in a fairly narrow window (something like 30 - 50 bits of actual information per second at average speaking speed).
If a language has fewer possible sounds, the syllables come more quickly. This is an innate tradeoff with a communication channel in information theory.
But the analogy is only partial because, unlike most designed protocols, the symbol distribution is not a normal or random distribution. And human language has, in information theory terms, internal redundancy and compression (pronouns do a table lookup) as well as error correction codes and checksums (most combinations are illegal).
This makes a analogy beteeen spoken language and a modem less than straightforward, though they are governed by the same underlying laws.
As in: to compare two languages A and B, translate large English corpus to A, then to B, then compare sizes. Ta-da.
And if actual speed is really what matters, feed results in some speech-to-text algorithm and clock the output.
I love the idea of parallelizing sentences to provide much needed context rather than having to reference the history in text or conversation.
…The paper found that, in terms of sheer number of syllables spoken per second, the fastest languages of the 17 studied were Japanese, Spanish, and Basque. The slowest were Cantonese, Vietnamese, and Thai.
…
What Pellegrino found is that, essentially, all languages convey information at roughly the same speed when all the factors are taken into account: around 39 bits per second. The higher the syllable-per-second rate, the lower the information density, which creates a trade-off that makes all languages around the same in terms of information rate.
So if you wanted to communicate the most quickly, you should rapidly speak in a more dense language like Cantonese or Vietnamese.[1] https://cheetahtemplate.org/
[3] https://www.cs.rochester.edu/u/scott/papers/1989_TR308_revis...