This is, I think, the best way to look at MT.
As many people here have commented, machine translations are often incorrect, misleading, or just incomprehensible garbage. However, in situations with enough context, people can use MT to accomplish tasks that they would not be able to do without it. Some examples, and a discussion of implications for foreign-language education, appear in a paper I wrote last year [0].
[0] https://researchmap.jp/multidatabases/multidatabase_contents...
Oh fascinating, do you have any idea of what things machine translation still consistently gets wrong?
Here's part of the Google translation of the front page of ja.wikipedia.org into English:
> 非行少年の再非行の抑止や更生を目的としており、決定までの過程として、「非行事実を家庭裁判所に送致・通告 - 家庭裁判所調査官(以下「調査官」と略称する)等による調査 - 調査結果をふまえた審判 - 必要に応じて保護的措置あるいは保護処分を決定」という流れを経るのが通例である。
> And for the purpose of deterrence and rehabilitation of re-delinquency of juvenile delinquents, as the process of until a decision, " the fact delinquency to the family court -up Reel-notice - family court investigators (referred to hereinafter as" investigators "), or the like by the survey - survey It is customary to go through the flow of " judgment based on the result- decide protective measures or protective measures as necessary ".
That's barely recognizable as English. The translation is hot garbage. Perhaps something you might use if you had an emergency.
Just to point out here... the text on the front page of Wikipedia tends to be very neutral, matter-of-fact, and well written. There are no characters (like in novels), no use of slang, and nothing else that should be tricky for translation, and yet we get this garbage.
It's a miracle that machine translation works at all, and when it does work, it is usually because you are translating between languages that are fairly similar to begin with. For example, English and Spanish.
"Its purpose is recidivism prevention and rehabilitation for youth offenders; the usual process leading up to a decision is: Notice or report of the fact of the youth crime to the family court > Investigation by the family court investigating officer (the "Investigator") and/or other parties > Judgment based on the results of the investigation > Protective measures or a disposition for rehabilitation, as necessary."
>Just to point out here... the text on the front page of Wikipedia tends to be very neutral, matter-of-fact, and well written
This is true in general, though perhaps not always of Japanese wikipedia, and certainly not on legal topics. I have no love lost with machine translation (hard agree on naniwaduni's comment about "insidious failure modes"), but there are very few humans who would be able to translate this particular excerpt effectively. Legal translation is my job and I still found this sentence to be reasonably challenging.
> The purpose of this process is to deter and rehabilitate the juvenile delinquent, and the process of making a decision on the delinquency is usually conducted in the following manner: referral and notification of the fact of the delinquency to the family court; investigation by an investigator of the family court (hereinafter referred to as "investigator"); trial based on the results of the investigation; and, if necessary, a decision on protective measures or protective measures. There is.
Translated with www.DeepL.com/Translator (free version)
DeepL is doing a better job of generating text that seems to make sense, but this just helps to more effectively communicate the wrong meaning, and in a manner less detectable to people who would use it. This is a much more insidious failure mode.
This is something I don't understand about the use of autocorrect. Autocorrect takes misspelled words, which are essentially always easy to understand, and replaces them with random other words. It's frequently impossible to guess what original word got mangled by autocorrect into appearing in a suddenly unintelligible sentence. This is a problem that misspellings just don't have.
Autocorrect is actively making communication worse. What's it supposed to be doing? Why is it there at all?
The first has your stated effect. The second does not.
Also, it has "bugs" that sometimes it will just drop an entire sentence or half-sentence for no reason.
"The juvenile protection procedure aims to deter and rehabilitate juvenile delinquents, and as a process up to the decision, "send and notify the facts of delinquency to the family court-family court investigator (hereinafter abbreviated as" investigator "). It is customary to go through the flow of "investigation by, etc.-a referee based on the investigation results-determine protective measures or protective measures as necessary"."
Not incredible (especially when it starts listing the process), but not as god awful as whatever you got. I'm curious if you omitted some part of the sentence when pasting in.
It's also worth noting that this is a pretty out-of-domain legal Japanese sentence.
Then there are other things besides pronouns that require inference. For example, adding な to the end of a sentence could mean either a modifier expressing admiration or confirmation, or it could be indicate a demand or order not to do something. In machine translation it almost always seems to be translated to "don't" no matter the context. That only makes the translation less accurate.
And I’d assume that could be the most common language people would want to translate from online platforms in this day and age.
Of course the details of how English is turned into incomprehensible garbage by machine translation always depend on the target language in question.
My biggest pet peeve with MT is the incredibly brain-dead choice of using English as an intermediate language.
This effectively prevents MT from ever becoming useful for any A→B translation where neither A nor B are English.
Example: put "пружи́на" into Google translate (or Bing translate, doesn't matter)and translate to German. You'll get "Frühling" as the suggested translation.
This is due to the fact that internally "пружи́на" is translated to English first, resulting in "spring". This is then translated to German with the statistically highest probability being "Frühling", even though "Feder" would be the only correct match ("Frühling" would be "весна" in Russian).
If you now say "But wait! This should clear itself up once you add some context, right?" you'd be wrong.
Try "Die Sprungfeder kam gerade rechtzeitig." ("The coil spring arrived just in time,") and you'll get "Весна пришла как раз вовремя." ("Springtime came just in time").
It's ridiculous!
EDIT: happens with DeepL as well; only Yandex gets it right (presumably because they actually use an actual German/Russian-model.
Do you have a source that that is actually what is happening behind the scenes? I work in this field and would be surprised if that is what Google is doing. My guess is this phenomenon is more due to a model that was trained on an Russian->English corpora first and then trained on Russian->German due to low resource, or something like that.
Someone like Yandex is going to have better translations because (shocker) they have larger russian corpora than google.
Indirectly yes, I do: https://codesachin.wordpress.com/2017/01/18/understanding-th...
> What this basically means is: If during training you provide it examples of English->Japanese & English->Korean translations, GNMT automatically does Japanese->Korean reasonably well! In fact, this is the biggest achievement of GNMT as a project.
It's exactly what Google have been doing for years now: train on X-to-English and English-to-Y only to get X-to-Y for free. Since the intermediate language only ever sees English as a source or target language, ambiguities like "spring" literally get lost in translation.
1. It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020.
> Since the intermediate language only ever sees English as a source or target language, ambiguities like "spring" literally get lost in translation.
2. This article is trying to dumb down what Google is doing - but you're right that this is why ambiguities get lost in translation, due to pre-training on a different language pair.
That said, there isn't literally a process of "translate into english" and then "translate english into german". These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German.
Sure, but that doesn't change the fact the training data is focused on English-to-X and X-to-English corpora. The underlying architecture of the model is just an implementation detail that doesn't really affect this as demonstrated by my example.
> These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German.
This is exactly what I'd argue isn't the case at all. Otherwise words that have a direct 1:1 translation wouldn't be mistranslated and companies like Yandex wouldn't be able to deliver so much better results.
German-Russian isn't low-resource at all, given 95M and 150M native speakers respectively and a close history for the past 150 years. [edit]The rich cultural history of both countries resulting in a vast library of literature, theatre plays, news publications, films and the general cultural relevance of both languages is even more important.[/edit] It's simply (quite comprehensible) bias towards English for research taking place in the USA and the fact that it's much easier to compile English-to-X and X-to-English corpora in a predominantly English-speaking country.
There are tons of translated books, films, news paper articles, scientific papers, etc. available for Russian-German and Yandex, being a Russian company, naturally has no problem compiling a Russian-German corpus (since they're not biased towards English).
This also happens between e.g. Spanish and Bulgarian.
For N languages you need to train N models if you go through English as an intermediate language, but N^2 if you don't use an intermediate language model. OK, computer time is cheap these days, but more importantly, for each language pair you need D high-quality documents translated in each of your two languages.
There are far fewer pieces of writing that exist in both German and Russian than in both German and English or both Russian and English.
There is cool-looking research on zero-shot language translation where the model learns its own internal "intermediate language". But you can be sure that if it produced reliably better results with the data available, the likes of Google would already be using it.
That's quite literally a cheap excuse. It'snot lack of training data (remember, you still need translations in the target language anyway!), it's a result of Anglo-Saxon cultural hegemony.
> There are far fewer pieces of writing that exist in both German and Russian than in both German and English or both Russian and English.
Right. And that's why Russia-based Yandex gets excellent results for German/Russian translation, I see...
> But you can be sure that if it produced reliably better results with the data available, the likes of Google would already be using it.
Nah, you're being way too optimistic there. Google is a business and as such, they don't chase SOTA outside of research (which doubles as advertisement). Good enough is the name of the game and it's simply infinitely cheaper to just train and maintain N-models (i.e. English/X and X/English) than to do the same for all combinations of X/Y-pairs (which would be N² models).
So as long as there is no incentive to do better, they simply won't, because they rely on the fact that English is the de-facto Lingua Franca and people translate to and from English more often than directly between other languages.
This would be an actual chance for companies from other countries (especially European ones) to step in and do exactly that, but again - "good enough" and reliance on English stops that from happening anytime soon. It's a missed opportunity especially with low-resource languages as demonstrated by the recent "Scottish" Wikipedia fiasco...
Definitely anything where a word has two or more different meanings.
Paper jam being my favourite example.
https://translate.google.com/#view=home&op=translate&sl=en&t...
It is translated as "(fruit) jam made out of paper"
As to sibling comment, IIRC Yandex translate has a nice feature where they'll provide a number of meanings for words when clicked upon in the source text.
"Invisible, Insane"