European word translator: an interactive map
ukdataexplorer.com
ukdataexplorer.com
p.s. I'm saying this because most of these terms that has a dial-map are common in daily conversation. The differences in written Chinese vocabulary aren't as significant; how scientific and technical terms are expressed is largely determined by your administrative region.
Edit: Ah, it says the data is from Google translator. So no suprise here, Google translator produces poor results. It's said that Deepl is much better. I can't really tell because I don't need machine translation Finnish English. Both are roughly equally strong foreign languages for me.
I think you cannot really compare the minuscule differences between "Standard German", "Austrian Standard German", and "Swiss Standard German" to the differences between English, Irish and Welsh, which are not even from the same language family. Also, the tool is based on Google Translate, and AFAIK Google Translate doesn't differentiate between them.
Comparing the tool to this map [0], it seems to do a pretty good job in capturing all major languages in Europe, while ignoring their dialects.
But I agree that I would be great if you could zoom into the map and also show differences in local dialects. ChatGPT seems to be pretty good at translating to different variants of standard German, or German dialects [1]
[0] https://en.wikipedia.org/wiki/Languages_of_Europe#/media/Fil...
[1] https://chatgpt.com/share/67bba4db-9458-800c-b5f8-fd3fa196d4...
- Swiss German and Austrian German didn't make the cut because Switzerland and Austria are on good terms with Germany and don't mind if we call their languages a dialect of German. Not only is that justification to exclude them, they are also not in Google translate for this reason (which this map uses)
- Luxembourg did mind and went to great lengths to get their German dialect recognized as a separate language, is in Google translate, but Wikipedia lists them as only 300k speakers
- Frisian is seen as a distinct language because of how different it is, is in Google translate, but has about 200k speakers
- Similarly, Scottish Garlic is in Google translate has only 70k-200k speakers
The map is consistent if you set the goal of only considering languages that are in Google translate and have at least 500k speakers.
I do think these rules detract from the map. Frisian and Luxembourgish are interesting as "in-between" languages (Luxemburgish has a lot of French influence, Frisian is closer related to English). And Swiss German has many distinct words that are very different from their German counterparts, so for the purposes of this map it really should be a language.
[0] https://en.wikipedia.org/wiki/Alemannic_German#/media/File:A...
But they can easily switch to more modest verion or even high german if needed.
This is absolutely not true. Bern is the capital and many people travel there for work or other reasons. It's also a dialect very heavily featured on TV (e.g. I remember there was a weather reporter from Bern, don't know if she still does this), a lot of famous politicians are/were from Bern (e.g. former Federal Council member Adolf Ogi) and many famous musicians also sang/sing in this dialect (Mani Matter, Züri West, Gölä, etc.)
Almost all Swiss dialects are mutually intelligible simply due to the high level of exposure to the diversity (and also their relative similarity). There are some people who don't understand Walliserdeutsch well, because it's less represented and also linguistically more removed from the rest - but even that's something you get used to quickly.
Differences between some of them are rather extreme, especially Prekmurje dialects feel like their own language - so we need to fallback to "book" Slovenian when talking with people from different regions.
[0] https://www.atlas-alltagssprache.de/wp-content/uploads/2014/...
For example in Rome a grocery store bag is "busta", but in Milan it is "sachetto" with "busta" being the word they use for an envelope.
one can follow migrations.. and criss-crosses..
btw, "orange" as color in Bulgarian is still "orange" (оранжев/а/о/и), but "orange" as fruit is портокал ("portokal") - so that's tricky..
"oranges" seems more correct, vs "orange color" maybe
Salt is an ancient Indo-European word that was already in use several millennia ago, so it has been inherited in most Indo-European languages.
Tea is a relatively recent borrowing in the European languages, which has spread from one language to another, with a few pronunciation variants, across all Europe, regardless of the genetic relationships between languages.
For instance "cow" and "Kuh" come from the same word as "boeuf" and "buey" (also despite the gender difference).
Arabic: mama babi. Mandarin: mama baba. Swahili: mama baba. Inuktitut: anaana ataata. English: mama papa. Tamil: amma appa.
These languages are not known to be related.
The first vowel sound a child makes is approximately "a" and the first consonant they form tends to be a nasal plosive "mba mba mba" and the second distinct sound tends to be a dental or labial plosive "pa ta pa ta". And the first thing a baby says is "mommy" of course and the second thing a baby says is "daddy" of course. So mama is mommy and papa or tata is daddy. That's the usual explanation, anyway.
That's interesting. Ana/ata means mother/father in Turkic languages
Here some other words as well: https://www.kleinersprachatlas.ch/karte-1-butter
countries receiving tea overland (e.g., via the Silk Road) adopted forms of “cha,” while those trading by sea through Fujian ports adopted forms of “te.”
The project visualise perfectly this distinction.
The term cha (茶) is “Sinitic,” meaning it is common to many varieties of Chinese dialects. Meanwhile, the word tea comes from the Min Nan variety of Chinese, spoken in the coastal Fujian province, where the character 茶 is pronounced te.
EDIT: playing with it, it's a bit sad that large numbers do not work at all (in any language); and that not all common forms of a word are shown. For example, I tried to see how "ninety six" is said in french in France, Belgium and Switzerland, but it does not work.
What is more influential (in a detrimental way) is German randomly switching reading direction. They read 2196 as 2000+100+6+90 instead of the more reasonable 2000+100+90+6
It takes a second to process and then they'll ask "do you mean [reverse order variant]?" so they do kinda get it and I think transitioning to the sane version could be possible without much trouble, but people would have to want to
five
six
seven
four
-teen
...and then cracks up. I cracked up when it was done to me, although apparently not everyone finds it so funny. ;)
I had schoolteachers who still spoke like this in 1970s Yorkshire. I don’t know if it was a regional dialect thing or a generational thing across all England, but among the over-40s back then it was still pretty common to hear German-style backwards numbers in English.
https://blogs.transparent.com/language-news/2016/08/29/danis...
1. Their words for Numbers are based on Base-10 system (so no nonsense such as „eleven” and „twelve”)
2. Their words for Numbers are short, one syllabe, so they can keep morę of them in their short term memory at once
Not saying there’s any truth to that, but sounds interesting
We never learned huitante (80), but here are apparently parts of Belgium that use is. We did learn soixante-dix and quatre-vingts-dix, and were allowed to use both. [0]
The Swiss also use huitante, and Nova Scotia uses octante.
[0]: Funnily enough, writing American English was a no-go. We had to write centre, colour, metre, lift (elevator), ticket (receipt).
Like how in Japanese, "mushroom" can roughly be translates as "tree child".
Edit: and a turtle is also a "shield-toad" (schildpad)
The words for bridge split neatly into language subfamilies. The only exception appears to be Welsh.
"Egy példa" would be more literal.
So it seems to not interpret the English word as the same one for all languages in case of words with different meanings!
She runs, as in the form of locomotion, is "ona biega/biegnie."
> This example demonstrates that the map should be interpreted with care; some translations have the meaning "she lasts" or "it works".
Another mistake for this example, although subtler, is the Dutch version, which is translated to the meaning of "she walks"
It's interesting that the site says it uses Google Translate, because using it via the web UI, it does give the correct answer.
https://translate.google.com/?sl=iw&tl=pl&text=she%20runs&op...
LLMs are so much better at this
(chatgpt 4o gave these answers:
English: folk
French: peuple
German: Volk
Dutch: volk
Italian: popolo
Danish: folk
)
Yeah, IME as well LLMs really shine at translation of sentences and getting the right meaning depending on the context for words. Way, way better than Google Translate via web UI or app.
I guess that right now there might be some high level IC/manager trying to get a promotion by switching Google Translate to use Gemini in a cheap and effective way :)
All such uses must be translated into different words in other languages. When the word to be translated has no context, a random translation choice is possible.
"Folk" as in "folk music" or "folk dances" may be translated correctly as "folkloric" or "popular".
The map would be more complete with this information because it may be very similar or completely different and can be interesting to compare, for example:
- EN: receptionist for both, NL: receptionist and receptioniste, DE: rezeptionist and empfangsdame. The map currently just shows the female version for German, without indication that they also use a transliteration of the English.
- EN: little brother, NL: broertje (the submission shows a doubled up version of kleine broertje), DE: kleiner Bruder. Although German has the diminutive suffix to make Brüderchen, they don't use it the way that we do, which I find interesting to see.
Google Translate's API can output multiple options, <https://cloud.google.com/translate/docs/reference/rest/v3bet...>, and Google's own website seems to indeed provide these different variants, but there is no label to say what the different array entries mean the way that Google's own website shows
I got curious which gender it guesses that you might mean. It seems to assume a male unless it's also very heavily female-connotated in English. In Dutch and German, it outputs male for hairdresser and doctor, female for nurse and receptionist (German translations mean "sick-sister" and "reception lady", respectively), and mixed for secretary (female in Dutch, male in German) because Dutch doesn't have a male word for it anymore (only workarounds)
Here are a few feedback for improvement,
- looking up for "user" gives some results, but starting from a term in other languages doesn’t work; also it doesn’t display multiple nouns that can apply as a translation (ex: ulisatrice/utilisateur in French)
- if the target user is English speaking user centered (at least for now), probably providing transliterations (along there Cyrillic/Greek correspondence) would probably make more sense
Input: cross
Russian: пересекать (as in verb "to cross")
Polish: Krzyż (noun, as a christian symbol).
Edit: "translation" -> "meaning".
At least they could have tried picking translations for matching parts of speech (nouns, verbs etc), and it would have been a great improvement, even if they ignored homonyms. Doing so does not require LLM.
Somehow it breaks on words "Monday", "January" and "Italy", for example - doesn't show any of the translations.
How does that work? Are they cached somewhere? Devtools doesn't indicate a call to google translate.
I assume it doesn't include any proper nouns. I tried putting in country names because I always find it interesting what different countries are called in different places, but it didn't return those either.
"girl" is my best so far
But of course, none of them are fully interchangeable in all contexts. You will typically not expect to hear "salut la compagnie" in a formal meeting with "les gens de la bonne société."
If you like synonyms, CRISPO gives 77 for société and 33 for compagnie.
Other than that, great implementation.
Or this village name: "Llanfairpwllgwyngyllgogerychwyrndrobwllllantysiliogogogoch", in it's short form just a smooth "Llanfairpwllgwyngyll" (https://en.wikipedia.org/wiki/Llanfairpwllgwyngyll).