Much of the Web Is Machine Translated: Insights from Multi-Way Parallelism
arxiv.org
arxiv.org
I can read three languages, and I'd like my browser to pick the "main" version of a page if it happens to be in any of those three languages, and only resort to translated ones if necessary. But there is no way to get this behaviour!
And I'd like so much for google to serve me Wikipedia pages in Italian or French if the subject concerns Italy or France, and only resort to English as a fallback, but no way, again!
If I set my main language to something other than English, I'll always get poorly translated MSDN documentation, and ugly web surfing in general. As such, I'm stuck with poorly translated youtube video titles, and English Wikipedia when searching about a monument in my hometown.
But <link rel="alternate" hreflang="..." href="..."/> tags are pretty common for SEO, so maybe an extension could parse those to check whether you would prefer one of the other versions.
No, it is not enough, because there is no way to tell: "Pick the source in my 3rd language rather than the translation in my 1st language".
For example, `.de` sites tend to be originally written in German, so I map the `*.de` pattern to the German language.
[1]: https://addons.mozilla.org/en-US/firefox/addon/accept-langua...
I don't think so. Accept-language doesn't allow you to differentiate between high quality and low quality content in the same language. With accept-language you could say that you prefer Italian or French over English where available, but that means you might also be served a crappy Italian/French machine translation on occasion, even though in that case you'd have preferred the English text.
(Microsoft for example seems to honour the accept-language header when looking at its documentation pages, but unfortunately that means that by default I get served the German page, which quite likely feels somewhat off due to having been machine-translated.)
In practice, however, it is difficult to obtain such a quantification. Further, if you could quantify the translation, you can also just improve them instead.
???
The q factor is just for ordering the language preferences from top to bottom (i.e. it doesn't matter whether I write "de,en;q=0.8" or "en;q=0.8,de" because both of them mean that I prefer German over English because German is implicitly q=1.0 and en is explicitly only q=0.8) and doesn't say anything about the quality of the content.
Applying this to Accept-Language, "de,en;q=0.8" means "I prefer German content unless it's less than 80% as good as the English alternative."
Of course in practice servers tend to treat all content they have as quality 1 and all content they don't have as quality 0 and then the lovingly crafted fractional quality values collapse to a simple ordered list. ("We serve what we have and you eat what you get.")
"12.4.2. Quality Values
The content negotiation fields defined by this specification use a common parameter, named "q" (case-insensitive), to assign a relative "weight" to the preference for that associated kind of content. This weight is referred to as a "quality value" (or "qvalue") because the same parameter name is often used within server configurations to assign a weight to the relative quality of the various representations that can be selected for a resource."
[0] https://www.rfc-editor.org/rfc/rfc9110.html#quality.values
[1] https://developer.mozilla.org/en-US/docs/Glossary/Quality_va...
Multi linguals face UX issues most professionals in the space are completely ignorant of.
https://addons.mozilla.org/en-US/firefox/addon/accept-langua...
By wonder, I don't necessarily mean it as in euphemism for "it is infuriating to me that X is Y" wonder: I've seen translation horror stories that, there allegedly are weird dogmatically motivated English translators that constantly abuse their power to just use foreign content as vehicles for their agendas, while original authors being unable to proofread themselves or chalking up noted oddities to their own skill issues until all is too late.
That kind of weird translation is not a typical experience for me relying translations to my primary language, and I suspect nor it is to most various non-English language speakers, but especially the recent wide and big push on MT as well as high praises on GPT translations seems to roughly in line with that abusive translator horror stories.
And here I wonder; is _that_ it, or am I overthinking it?
I've seen references to this as well, but I've not seen a source that isn't themselves trying to stir up some sort of drama or bait a particular fandom. Could you elaborate?
I suspect the underlying phenomenon is simply English monolingualism - educated people around the world are generally expected to understand at least one second language, which is usually English due to the Internet and Hollywood. So they're familiar with being able to read both sides of a translation and judge it. While the monoglots simply regard the production of a piece of text they can't read by a machine as job done.
If my Accept-Language header is "no-nb, en-us, en-gb" it means that I have declared that I _know_ all these languages, but if you have a high quality version of Norwegian I'd prefer it. Do not try to be "helpful" by forcing a machine translation from English to Norwegian on me. However if there is a useful review in French, providing it machine translated is perfectly fine.
All Google products are notorious for this. Google Play, Google Maps Reviews etc. will be poorly translated into Norwegian, instead of being left as is.
Croatia is a small country and we don't have vast amounts of content in our own language, so when I search I often get bombarded with machine translated content from Russian origin. I suppose it is Russian, since there are images with text written on them that is in cyrillic.
If it were Serbian (they also use cyrillic) there would be no need for translation since we basically speak the same language and a lot of Serbian content is in latin alphabet also.
This has some truly weird implications, e.g. once language model translations switch from subtitles to dubbing it will effectively stop shifts in pronunciation, because the models won't be re-trained with an ear to the street.
This has already happened to a great degree thanks first to radio and the television, as well as the early 20th century movement to “received pronounciation” in many countries. E.g. TV finally killed off thee and thou in the 1960s
It often feels like there are regional accent variations but they are usually quite minimal these days thanks to spread of technology.
The in Europe, explicit suppression or uniformation of language began AFAIK with Louis XIII and his deliberate formation of “France” (as opposed to just a collection of regions controlled by one person). In China I believe the same thing was instigated by the (by coincidence contemporary!) Qing dynasty, but it might have been a lot earlier. In any case it really zoomed throughout the world in the 20th century when communication technologies and practices were adopted by the emerging nationalist movements.
So machine translation will simply continue a longstanding process.
Rapid communication over long distances in time might be able to put a stop to that, e.g. if teens end up interacting more with simulacra of long-dead actors than others of their own age, there could be some weird effects. But I think that's unlikely to happen, since someone is bound to come up with a more popular version that has all the newest slang.
However now you mention it I wonder if automatic translation might also cramp or otherwise affect minority languages spoken by a small population.
Also I wonder if multi-step automatic translation will cause weirdness in spoken languages (again, mainly in minority languages)
Quoting this para because it's so good as a statement of requirements:
“What we need today is a readable, audible, singable, speakable, dictatable language which we can read aloud without the need to translate into the spoken language, with the help of which we can take notes without the need to translate into the literary language, which we can [use] at the speaker’s desk as well as on the stage, and which even village grannies, women and children can understand if we read it to them. Any language that does not meet these requirements is not a living language, and can under no circumstances become the national language of our country.”
Can someone provide a source for this? Could make for some interesting discussion points, but I don't want to propagate unverified information
Losing that accent was deliberate, if sad, but IME Americans aren't really tolerant of non-US accents, even when they find them cute. And the speech recognition systems are definitely intolerant.
The decisions about using Python vs JavaScript vs Go vs .NET vs Rust vs Java etc etc are going to quickly become moot when businesses realize they can get value without focusing on any single tech stack. The next logical step after that is reducing the inefficiencies of having so many tech stacks to maintain - I would guess by having more and more programming languages fall to the wayside, probably starting with the ones that aren't popular and ending with the ones that aren't easily ML-usable.
The benefits of choosing a common language are obvious and historically tested, as already pointed out in parent comment. I’ll add that today, English is the language of business. Many people have learned English in addition to their native tongues. Not because English is so wonderful, but because it gains benefits in the real world of commerce, education, politics, etc. This is a familiar concept historically. There’s a reason for the phrase Lingua franca, and for Latin being the language of the medieval church.
Now, why are there a multitude of programming languages? I know, the right tool for the right job. I say that all the time too, but that is missing the point. For computer programming, a multitude of languages was really the only solution for a long time. Now we have a prospect of a human taking a concept, telling AI, and having AI provide what is necessary for the computer to do the needed work. We’re in the midst, just like for a time a culture might learn several languages, in addition to their own, in a trade route. But the trends over time seem obvious to me.
The extent to which US culture war topics drive UK politics infuriates me. Everyone wants to be part of the bigger, flashier show rather than deal with real things.
(Big exception: Japan. China could have gone this route but more or less has chosen not to because of the internal political need to suppress its creative industries)
It wouldn't surprise me if there were some great translations of great Bulgarian content that are simply not as prominent among the flood of other content available in English.
I just skip almost all our content now and go directly to english speaking sources. Press, TV, radio, online articles, books, yt videos -- a big part of them is more or less "inspired" or straight up licensed. Very difficult to come across truly original content, even when it's claimed truly original (e.g. dancers like to tell that their shows are original, but they take "master classes" from foreign teachers anyway).
I bet you a lot of people even like it and would defend it. That could be why it is prolific.
My idealist bubble is often burst when I'm out actually interacting with them.
Really annoying and limiting.