Also there was a Swedish word that was just copied ("anullerades", meaning more or less "cancelled", "voided" or something along those lines) into the English text.
But, overall, it's still rather impressive.
Also there was a Swedish word that was just copied ("anullerades", meaning more or less "cancelled", "voided" or something along those lines) into the English text.
But, overall, it's still rather impressive.
For example, the word "Eesti" will often get Google-Translated to "English" rather than "Estonian".
This means that a film at my local cinema that my web browser assures me is in "English" will in fact be in Estonian. And an interview with a Russian saying that he doesn't speak Estonian gets translated so that he appears to say that he doesn't speak English.
The product designers special-cased language names, doing extra work to produce what will almost always be the wrong result.
(And what they can do to place names is often patently ridiculous. For example, "Peterburi tee" should either be left alone or maybe translated to "St Petersburg Road" but actually somehow becomes "Hertford Road". And the ZIP + City name "13415 Tallinn" becomes "thirteen thousand four hundred and fifteen Tallinn".)
They absolutely did not do this. It's an artefact of statistical translation. In the corpus there are a lot of English documents saying "This document is in English", whose translated versions in Afrikaans (because I know Afrikaans) say "Hierdie dokument is in Afrikaans". Thus the translator learns the "hierdie" is Afrikaans for "this", "dokument" is "document, ..., and "English" is "Afrikaans".
The street name issue probably comes from an organisation whose Estonian office is in Peterburi tee and whose English office is in Hertford Road.
Are you sure about that? It would sound quite likely to me that simply, the word for "English" tends to appear in the same context (N-gram etc.) as the word for "Estonian". For example the sentence "I speak English" would be common in English, while "I speak Estonian" would be common in Estonian, so it might associate the words together.
Anyway, plenty of text would probably match word-wise ("X bought Y for $AMOUNT $UNIT") in financial news, so the mapping of "dollar" over "kronor" seems a reasonable error.
For probably the same reason, google also translates hungarian "1000 forint" to english "1000 HUF" going from the full word to a quasi-acronym for "HUngarian Forint".
I saw some article once, where the names of a prime minister or some such was "translated" to the name of the US president. Weird.