This makes me wonder, what is the largest difference between letter count in two different languages?
This example has a 4:19 ratio. Depending on what translation you go with (I think the consensus is actually the three word answer “nimium saepe valedixit”), the Latin example has a 22:38 (11:19) ratio.
Of course, this is just considering alphabetic languages. If we look at SE Asian languages we will find more extreme examples. For instance, a google search led me to:
“If we're going the other way, it could be "伥", which Pleco gives as
the ghost of a man who fell a victim to a tiger, yet helps the tiger to devour others”
https://www.reddit.com/r/ChineseLanguage/comments/6ijiuw/lon...
Nimium saepe valedixit is 9 syllables and, as frequently noted on the page, does not attempt to translate the entire English source text, which is 10 syllables. It was kind of surreal reading the answers, since none of them attempt to determine what the English lyric means, and it can't be considered fluent English when seen as an isolated sentence. You need to determine what it means before you try to translate it into another language.
I just listened to the song (well, the first three verses, which is all of the verses) while looking at a printout of the lyrics, and I can't determine what that line in the chorus is supposed to mean. It's very strange grammar:
This love has taken its toll on me
She said goodbye too many times before
Her heart is breaking in front of me
And I have no choice
'Cause I won't say goodbye anymore
The line in question, She said goodbye too many times before, stands out like a sore thumb for being preceded and followed by sentences that, unlike it, are both in the present tense. There is no indication anywhere in the song, as far as I can see, of what "before" refers to.
So my instinct is to essentially write off the possibility of translating the lyric with the aphorism "garbage in, garbage out".
Is English your native language? I’m asking because it is my native language and the entire phrase, including “before”, is clear to me.
“Before” is a temporal indicator. You could replace it with another temporal indicator and the phrase would still make sense, For example, “She said goodbye too many times today”. You wouldn’t ask what the antecedent is for “today”. Same with “before”.
> I’m asking because it is my native language and the entire phrase, including “before”, is clear to me.
It is a common phenomenon for people to claim that sentences are perfectly clear to them when, objectively, those sentences do not have a meaning at all. On Language Log they occasionally discuss "Escher sentences", with the prototype example being "More people have been to France than I have".
> “Before” is a temporal indicator. You could replace it with another temporal indicator and the phrase would still make sense, For example, “She said goodbye too many times today”. You wouldn’t ask what the antecedent is for “today”. Same with “before”.
Except I can see what's happening with "She said goodbye too many times today." That sentence will be followed up with some explanation of the consequences of having said goodbye too many times.
In the chorus, the intent might have been that the line "she said goodbye too many times before" is an explanation of the preceding line (that's how people are interpreting it here). Or the line might just have been thrown in with no rhyme or reason, completely disconnected from the rest of the song. But regardless of the intent, the line has failed to connect to the sentence before it or the sentence after it, which means that we cannot determine what it's trying to say.
> You wouldn’t ask what the antecedent is for “today”. Same with “before”.
Moving back to this, it's necessary to ask what exactly "before" is referring to because the question came up of whether and how it should be represented in the Latin translation. It might conceivably refer to "before now" (in which case the suggestion of Latin perfect tense is fine), "before some point identified by the context" (you'd want pluperfect, if the point was in the past, or future perfect if the point was in the future [or of course perfect if the point is "now"]), or "before some specific event" (you'd want the preposition ante, and you'd also need to mention the event).
That is a strong possibility. It doesn't solve the problem with the line; to make sense, it should say she's said goodbye too many times before.
> It sounds clear to me: the girl has previously tried to break up many times, so now he is breaking up with her.
That is not so strong; the first verse is phrased in a way that suggests she is leaving him, not the other way around:
I was so high, I did not recognize
The fire burning in her eyes
The chaos that controlled my mind
Whispered goodbye as she got on a plane
Never to return again but always in my heart, oh
(On first impression, I assumed this verse meant that the girl was dead, but she could just be leaving.)
And the first line is past tense, just like the second.
Edit: Reading all of the lyrics they were sleeping together, she fell in love with him so he broke it off
This is a somewhat complex issue, so please bear with me.
First, we can dispense with the idea that the tense of the first line is "just like the second". They are different and the difference is quite significant.
Whether the first line should be called "past tense" or "present tense" is more of a fussy terminological issue. There are two concepts in linguistics which have to do with how the verb relates to a timeline:
- "Tense" has to do with whether the action takes place before, during, or after whatever time would be referred to by the word "now".
- "Aspect" has to do with the temporal structure of the action itself, rather than its position relative to a "camera" placed at "now": maybe the action occurs at an indivisible point in time ("That's when I noticed the rabbit"); maybe it takes place continuously over an extended duration ("I've been reading for thirty minutes"); maybe it occurs at a large number of separate points within a continuous window ("I used to visit the donut shop every day after school")
Except I used the wrong words just now. "Tense" and "aspect" are terms from syntax, and you can determine them purely by looking at the form of the verb. The definitions I gave belong to semantics: when I said "tense", I should have said "time", and I'm not sure what the semantics-specific term for the quality related to aspect is. Anyway, we name the verb forms, "tense" and "aspect", according to whether they primarily correspond with those semantic definitions.
Except, again, there's a little more to it. We'd like to name the verb forms according to this distinction, but there is a long tradition in Latin scholarship of referring to both of those distinctions by the same name, "tense", and this bled over into English.
So we can say the following about line 1 and line 2:
- Line 1 is, semantically, focused on the present. It is making a claim about "now".
- The verb is conjugated in what would traditionally be called the "perfect tense"; according to the tense/aspect distinction described above, it is present tense (reflected in the form of have), indicating that we are talking about "now", and perfect aspect (reflected in the fact that have is used at all), indicating that the action described ("taking a toll") is already finished.
- Line 2 is semantically focused on the past. It is making a claim about some time before "now".
- The verb in line 2 is conjugated in what would traditionally be called the "simple past" or "preterite" tense. The aspect is not clear, because the English preterite tense is used for multiple different verbal semantic aspects.
The fact that line 2 is talking about the past when the rest of the chorus is talking about the present is very strange.
Can't it? Why not? What's wrong with it?
> The line in question, She said goodbye too many times before, stands out like a sore thumb for being preceded and followed by sentences that, unlike it, are both in the present tense.
But it's a song. Prosody can't be held to the same strict rules of tense consistency (or other grammatical rules) as prose. And flipping tenses between lines is hardly an uncommon feature of songwriting. Take, for example, Leonard Cohen's "Boogie Street":
A sip of wine, a cigarette,
And then it's time to go,
I tidied up the kitchenette,
I tuned the old Banjo.
I'm wanted at the traffic jam
and so on.
She started her new life
Ten dollars in debt
That's all it took to get started back then
A trip to the courthouse across the state line
No one could stop her
She'd made up her mind
He was eighteen
And she wasn't
But she said she was / and never thought twice
And came back home as my daddy's wife
She just shook her head
When her mama said "Are you sure he's the one?"
But she was
Here we see some fairly complex temporal structure handled fluently, with no problems of any kind. The writing is better.
>> it can't be considered fluent English when seen as an isolated sentence
> Can't it? Why not? What's wrong with it?
The use of the simple past tense is not compatible with the sense of before that everyone here is trying to assign.
He's lamenting a romance that's been difficult for him. Before now, during the difficult relationship, she said goodbye or left him too many times, causing the difficulty and toll it has taken on him.
Other translations are:
So it happened in the course of time
So it came about in the course of time
And in the process of time it came to pass
And it cometh to pass at the end of days
If the accepted answer stands, that's remarkable. I wonder how one could measure a language efficiency. Maybe syllable count ? But one would need a sort of assembly to translate to and verify that a sentence computes the intended information.
“Darmok and Jalad at Tanagra” begs to differ
then don't these other langs do a 'bad job' at compression ?
> That doesn't seem feasible
In general maybe not. But for some restricted 'assembly' ?
"The cat is on the table" has no ambiguity. And in some langs like Polish it compresses better : "Kot jest na stole" (Cat is on table), same info, better syllable-wise compression (5 vs 7).
For instance, your assumption that the definite "the cat" is being used idiomatically like so: this sentence, used in the manner you offer, might be used in conversation might occur in a farmhouse somewhere between an old man and woman who have lived together in this house for a long time, i.e. American Gothic. There's a vast amount of shared information and a perception of very little ambiguity held by both the speaker and the listener (whether correct or mistaken!). Any of those might fail. Furthermore, to use this sentence in English unadorned by context requires that both the speaker and listener have a shared reference to _what_ cat is being referred to by the definite article, "the". This very well might come with an unambiguous default in other languages!
Translation only gets more complicated from this.
Kanji look very compressed, and can convey a lot in a single character, but if there isn't one for your needs things can get ugly. Whereas English speakers find it much easier to borrow, shorten, abbreviate, or make up words for convenience.
In Turkish if you say kedi you mean the cat. The copula is in the third person is conveyed by an empty word at the end of a sentence.
Masada bir kedi var.
There is a cat on the table. A cat is on the table. A cat is said as one cat = bir kedi.
*Bir kedi masada.
Doesn’t work. This says a cat on the table. It’s an ellipsis.
However, traditional way of computing efficiency of compression would not be useful for a meaningful analysis of the efficiency of a language. Barring issues like having an ideal encoding to bits, or even having the concept of "efficiency" being rigorously defined, there are problems just from the outset.
Take context for example.
All useful compression methods have some sort of decompression key involved. This could be the dictionary, or the bitmap or the know-how (for cases like RLE). In natural langauges, the compression/decompression key is stored in a distributed fashion across the minds of a society.
"Darmok and Jalad at Tanagra" is a VERY efficient compression for what is presumably a very long story about two hunters who met at an island and fought a beast together, but it is only efficient to the people who speak that language. The "local" efficiency (to the population who speak the language) is very high, but the "global" efficiency isn't.
So we must account for efficiency in terms of the size of the compressed concept as well as the compression key. And from my experience, it's a sorta lumpy kinda world out there.
This has been done! The answer is about 39 bits a second.
https://www.science.org/content/article/human-speech-may-hav...
ISTM a similar principle would need to apply here: learning the "Darmok and Jalad at Tanagra" language would involve absorbing many volumes of history and mythology where for the usual sort of language a dictionary, grammar reference and maybe a book of common idioms would suffice.
Whatever metric is used to compare languages for efficiency should reflect this.
I suppose image macros/memes are the modern equivalent. Social context enables readers to "decompress" the meme.
[Drake top]: "Two hunters who met at an island and fought a beast together"
[Drake bottom]: "Darmok and Jalad at Tanagra"
https://www.science.org/content/article/human-speech-may-hav...
An area of active research:
* https://www.theatlantic.com/international/archive/2016/06/co...
* https://www.degruyter.com/document/doi/10.1515/lingvan-2020-...
A constructed language:
* https://en.wikipedia.org/wiki/Ithkuil
* Via: https://www.reddit.com/r/linguistics/comments/zyi8q/what_is_...
Doing a search for "language information entropy" also gives back a number of results.
Similar thing I’ve noticed with the south indian language - Malayalam, just try to pronounce the name of the city - Thiruvananthapuram, local speakers would pronounce it with roughly the same speed as “London”, and would enunciate every syllable - its crazy.
No magic but plain agglutination, and I am sure this should be possible in languages like Finnish too...
For example, mandarin or japanese can be very short on the character count. However, this increases character complexity and makes the languages harder to learn. On the other hand, large parts of english tend to be simple to learn.
Translation is a process which both erases information and introduces new information. Any comparison of languages which tries to evaluate which languages are more compact has to work with some assumptions about what information should be conveyed. A statistical distribution of language-independent messages. But when you choose a distribution, you’re encoding your biases.
Not saying that language efficiency is a bunk concept, just that it’s a thorny, difficult concept to quantify. Same is true of data compression algorithms—there is no such thing as an absolute scale for Kolmogorov complexity, for the same reasons.
That’s why the translations of Asterix are so impressive.
There's an essay I enjoyed by Douglas Hofstadter which is all about this, though from an artistic POV without much (any?) information science. Translator, Trader. The title itself is a fun bit of translational wordplay on "traduttore, traditore."
"Translation is like a woman. If it is beautiful, it is not faithful. If it is faithful, it is most certainly not beautiful."
When translating for fun, I've often run into a choice between:
- Preserving the author's meaning as literally as possible.
- Preserving the author's style.
A translation can be literally very accurate, while destroying everything that made the original work charming. Or it might preserve the feel and the flavor of the original work, but skim over a lot of the details. A really good translation captures more of both, with fewer trade-offs.
Jorge Luis Borges encouraged his translators to improve upon his original work, if possible. He worked extensively with Di Giovanni, one of his translators, debating the best way to capture certain phrases in English: https://medium.com/@michael.marcus/dear-mr-borges-which-tran... His preference was almost always to capture the "feel" of the work, rather than a strictly literal translation.
I have an odd book, which contains three copies of the same story: An original in English, a French translation, and then a translation back into English by a new translator. The French version definitely loses something, and the second English version loses a bit more. But in the second English version, there is an occasional delightful turn of phrase, something that's briefly better than the original version. Translation is hard.
https://en.wikipedia.org/wiki/Breathless_(1960_film)#Closing...
I'm not sure if the Latin translation as it is captures that, or if you'd go for something more like "valedīxisse nimium valedīxit", "she said goodbye having said goodbye too much"; I kind of like that because then you're saying goodbye twice in the same line.
Exactly (or at the time). Many submissions here attain shortness by eliding this important precision.
In an optimally breve language the meaning of a text could completely flip when a single letter/syllable/phoneme is changed. That, in turn, means listeners have to hear every letter/syllable/phoneme perfectly.
Interestingly, natural languages already have a bit of both.
As an example, if you skim-read a text and restart at “He said she wasn’t there anymore”, there are 3 ‘back references’ in that sentence that require you to look back in the text to find the meaning of.
Also, a paragraph’s meaning can change by adding the sentence “Just joking.” Or even a simple “Not.”.
"nimium valedīxit": He got too sick
"totiēns valedīxit": He was always well
Edit: Playing around with google translate, "nim valedīxit" translates to He said goodbye. But "valedīxit" translates to Said goodbye. "Nimium" translates to Too many
So somewhere in that complexity it does seem to be that those two words have a meaning that build off eachother for their meaning, but google is considering it literally
If anyone has an explanation for these phrases rather than my guess work, I'd love to hear them!
totiēns valedīxit: She said goodbye so many times.
nimium valedīxit: She said goodbye too much
https://chat.openai.com/share/6d564b0a-c613-4411-a656-735cd9...
Because while "classical Latin" was capable of doing those antics, it was limited for day to day use. Phrasal and noun endings were complicated and wouldn't play well with day to day usage
Finnish doesn’t have gender pronouns so you can’t distinguish between he and she in most contexts. Adding that distinction in an idiomatic way would make the translation quite a bit longer.
The ancient languages like Old Arabic, Old Hebrew, and Latin was the key to understanding language in general. I think Esperante might also be key to deducing language.
Eng: "I dare you to drink that"
Deu: "Ich fordere dich heraus, das zu trinken."
Almost doubleAnd we have more of those, like "schon", "gell", "fei" (in some dialects), "halt", "eben". Maybe a few more I cannot think of right now.
Do you have any examples where it really excels? In my experience English is quite a good language to describe complicated things rather simple and short.
> So I would cut this down to something like nimium valedīxit or totiēns valedīxit: "she bade farewell too much before" or "she bade farewell so many times before".
Édit: I mean in the last paragraph of the answer.
> nimium valedīxit or totiēns valedīxit: "she bade farewell too much before" or "she bade farewell so many times before".
nimium/totiēns conveying "too many times before" and "valedīxit" conveying "she said goodbye".