DeepL Translator – AI Assistance for Language
deepl.com
deepl.com
DeepL performs significantly better on the most difficult texts I've given it. It's substantially better with colloquial language, and—oddly—nautical language. It also seems to be much better at tracking relationships between words in longer sentences. Good work!
Original: Angela Merkel hat sich gegen Vorwürfe gewehrt, dass es dem Bundestagswahlkampf an Spannung fehle.
Google translate: Angela Merkel has reproached himself against allegations that the Bundestag election campaign is lacking in tension.
DeepL: Angela Merkel resisted accusations that the Bundestag election campaign lacked tension.
English again: Angela Merkel opposed the accusation that the election campaign in the Bundestag was not tense.
German: Angela Merkel wandte sich gegen den Vorwurf, der Wahlkampf im Bundestag sei nicht gespannt.
English: Angela Merkel objected to the accusation that the election campaign in the Bundestag was not tense.
German: Angela Merkel wandte sich gegen den Vorwurf ein, der Wahlkampf im Bundestag sei nicht angespannt.
This is a fixed point (the translations no longer change).
The quality and stability of the translations is impressive, but the final German is a bit off, it seems to be confused between "objected to [something objectionable]" and "objected that [some counterargument]", mixing both in the same sentence.
EDIT: I tried going through all languages, German->English->French->Spanish->Italian->Dutch->Polish->..., after a few iterations, it settled on "Merkel wendet sich gegen die Vorwuerfe, der Bundestagswahlkampf sei nicht gespannt." (Merkel is opposed to accusations that the Bundestag election campaign is not tense.)
Things that got lost in translation: Merkel's first name and the past tense (EDIT: and the subtle distinction between "not lacking tension" and "being tense"). The pluralization of accusation(s) seems to change based on the language. Really quite impressive to maintain the meaning over so many steps.
And by the way. It's impressive^W^W Its impressive handling of the abomination that is "dass es" and similar constructions is pretty impressive (see what I did there?). In that sense, I'm surprised the reflexive form was recovered in the German.
* "The impact of the solar wind protons on the surface of Mercury" became "Der Einfluss der solaren Windprotonen auf die Oberfläche von Quecksilber". Note that 'solar wind protons' should have been translated as 'Sonnenwindprotonen' instead, i.e. the word 'solar' was to be a part of the noun's modifier, but it was pushed out.
* The lack of domain-specific training is especially obvious with the case of the planet's name being translated as "Quecksilber" instead of "Merkur" (Quecksilber being the name of the metal).
* "pure northward interplanetary magnetic field (IMF)" became "reines interplanetares interplanetares Magnetfeld nach Norden (IWF)". Aside from this being a poor translation, it's worth noting that DeepL didn't properly process the introduction of an abbreviation (IWF being the abbrev. for the International Monetary Fund in German).
Google: La fruta vuela como una flecha.
DeepL: La fruta vuela como una flecha.
"Fruit flies like bananas".
Google: La fruta vuela como plátanos.
DeepL: Las moscas de la fruta son como los plátanos.
Even so, the correct translation for the second sentence would be one of "Moscas de la fruta como plátanos" or most probably "A las moscas de la fruta les gustan los platanos" instead, the ambiguity is due to like being either a verb or adverb.
Way to go guys!
"I saw some weird fruit flies." "What were they like?" "They were like bananas." "Fruit flies like bananas?" "Yeah, they were implausibly yellow and banana-shaped."
By the way, most machine translation models are trained on news data. Try some out of domain data like tweets or other social media comments, if you want to put it to the test.
https://www.golem.de/news/deepl-im-hands-on-neues-tool-ueber...
http://www.lastampa.it/2017/08/29/tecnologia/news/deepl-trad...
Google Translate:
> Better than Google and Microsoft - the German company DeepL is committed to translation services. DeepL uses a novel architecture of neural networks and uses a supercomputer with 5.1 petaflops. By the same company, the service comes Linguee who has already made with translations of individual words or phrases a name. Texts translated by humans are used for translations in order to provide better results.
DeepL:
> Better than Google and Microsoft - this is the goal that DeepL, a German company, has set itself, at least for translation services. DeepL uses a novel architecture of neural networks and relies on a supercomputer with 5.1 petaflops. The Linguee service, which has already made a name for itself with translations of individual words or groups of words, comes from the same company. In doing so, human-translated texts are used for translations in order to deliver better results.
I think the DeepL version makes more sense here, but I am not a German reader so I might be off.
(Unfortunately EN->DE translations from DeepL are not on the same level - yet -, at least not on the stuff I just put in.)
"And if thou wilt not, I shall need violence."
Kinda sad to hear, but completely understandable. I'm curious whether the difference in performance is due to their model specifics or just better training data.
Does anyone have more information?
I have no information on the model, unfortunately.
"Twas Brillig, et les fentes fendues tournoyaient et gimblaient dans l'épée. Tous les mimsy étaient des borogoves, et les mome raths dépassaient les ragots.
Méfie-toi du Jabberwock, mon fils! Les mâchoires qui mordent, les griffes qui attrapent. Et'ware l'oiseau Jubjub, et fuyez le bandersnatch frumieux.
It's curious that it didn't understand "'twas" for the French translation, but apparently did for the German one. (My German is almost nonexistent, though.)
I've never been able to find one, but maybe I just haven't looked hard enough.
I figured someone would have gone to the trouble of combining models with a maintained collection of datasets to produce an open source alternative to Google Translate by now. I've been wondering that for years and it never seems to happen. Not saying anyone should feel obligated - I'm just curious why we don't see this, when we see so many other open source software projects that are competetive with their commercial alternatives.
Is it difficult/expensive to acquire these datasets? Is it a lot of effort to actually fine-tune the algorithms to reach passable results?
It seems (without knowing the details myself) that the state of the art in actually usable machine translation tools is always locked up in commercial IP, even though it feels (at least to me) like something that should be a free public service and therefore an ideal candidate for the 'open source' treatment.
Über allen Gipfeln
Ist Ruh,
---
Above all summits
Is Roo,
I wonder if you could find more artifacts like this to find the training data they used.