Has Google Translate been fixed yet?
isgooglefixed.tw
isgooglefixed.tw
- mega → 巨型 (百萬)
Really? Do you really want to translate mega to 百萬 (million) in every context?
*Edit: every, not any
- text → 文本 (文字)
Seriously? 文本 is the formal, academic way to say it, even in Taiwan. Have this person never been any academic environment?
- through → 通過 (透過)
Seriously???
Would its use as a metric prefix not be a correct context for this?
*Edit: my word choice was wrong in the parent comment. I meant to say "every context", not "any context".
The initial comment looked like there wasn't any context, which means "mega" could just as easily mean «SI million prefix» as «huge».
In written English, the word "minute" in isolation and without context may be the noun for the time period of 60 seconds, or it may be the adjective «very small».
This particular usage has been around, colloquially, for a minute.
In any case, Bing and DeepL both agree with Google on translating it to 巨型 (giant).
When I was a kid, my best friend was chinese, and we could never do anything on saturdays because he had to go to a school to learn chines, and how to write in chinese was very hard and time-consuming.
Do you know? Care to explain?
EDIT: I am not from the US, but from Argentina.
EDIT2: I am talking about the effort to learn how to draw (not only read but draw!) very complex looking 3000 icons vs. the effort of learning 24/26 and some rules. With a western language, by just knowing a few pronunciation rules, you can read any word. At least in Spanish. I am not implying that Chinese are bad, or that the language should die or anything like that.
(Also, learning to write English probably consumed more of your time than just a Saturday per week.)
I'm not sure if you're aware, but asking why an ethnos/nation hasn't significantly changed their language/culture to appease someone who finds it difficult to grasp comes across as incredibly rude.
I found that believing that a person trying to understand something and asking a respectful question is rude is incredibly rude and ignorant.
Neither is objectively superior to the other, you are just used to what you were born into
I do want to stress that any language that is old enough will eventually contort into a state where the writing is lagging behind the spoken language as it develops faster, and eventually lots of room for optimization starts to appear, lots of legacy to remove. As the culture develops and the spoken language simplifies and words get added, ambiguity seeps in, rudimentary language construcs start to appear unfamiliar to the commmon speaker. The vocabulary now consists of a large mix of old and new, with some redundancy and barbarisms thrown in for good measure. And now, there's lots of room for improvement. And sometimes a nation (or its authorities) decides that its time to simplify things, as mainland China and Sweden and others have done. I wish someone did this with English, bit since there's a multiplicity of English speaking countries, there will never be a meaningful overhaul that doesn't turn into a massive mess.
If he juse had used Spanish, he would have been learning piano or something on Saturday.
When I read anything, be it mandarin or Japanese or English, I attach meanings to words first. In fact, I am attaching meaning to logical structures and phrases, and then the individual words make the detail. Converting words into sounds seems to be a different skill from converting words into meaning. It really doesn’t matter whether the words are made of strokes or letters of the alphabet, the breakdown of the little details is a separate mechanism from comprehension.
Well, I still think Latin alphabet (or the Arabic one, but I don't know it enough to be sure about it) is more optimized / easier to grasp conceptually when you don't know the word. For example one of the compounds illustrated mixes 2 pictographs, but one is used for the meaning and the other for the sound of the final word (the "to wash oneself" example); that doesn't sound easy to me.
For reference, Chinese people think their language is very easy and that English is absurdly hard (for example, that we have an absolutely unnecessary yet mandatory number of ways to indicate tenses and plurality, and if you screw them up we think you're stupid even though there's no legitimate functional use for the differences - to say nothing of the fact that even with simple pluralisation like adding an "s", that's not one sound, that's three distinct sounds that a Chinese person would need to learn). A nationalistic Chinese on Weibo might air the opinion that the Chinese script is more optimised/easier to grasp conceptually when you aren't familiar with the concept than English, where you have to have familiarity with the suffixes and prefixes of two separate non-English languages (one of them dead!) as table stakes, plus a lot of familiarity with French. Of course, this wouldn't be correct, most people don't really understand a new word they come across in that much detail (in English or Chinese), it's just because you're not a speaker of the language and it's in a different script that it seems so hard to you. I promise you, Chinese is not objectively harder than English to learn. Chinese children have the same language development timeline as English and Arabic speakers, and Chinese scientists are very prolific and accomplished. Chinese is just quite difficult to learn for an English speaker.
The simplification did make it easier to write by hand (which is becoming less and less relevant with computers) but doesn't necessarily make it easier to learn new characters because of this.
The current system works, replacing it would be a massive multi-decade long change with a lot of opposition. Cultural pride/arrogance plays a big role. It's not inconceivable to think that the Chinese regime might fall in the process if it tried to pull it off. Given all these, the benefits just don't outweigh the costs/risks, at least in the horizon of the next 100 years.
So a different writing system for Chinese has already been developed.
Some other languages also have more than one writing system. (Serbian and Azeri).
When mobile phones became accessible to the general public then many people used Latin alphabet to write in Russian, because SMS didn't support Cyrillic properly.
Vietnamese also didn't use Latin originally.
Tangent: Anthropic's CEO worked on Mandarin speech recognition at Baidu: https://arxiv.org/abs/1512.02595
Quiet environments were, for me, more of an issue: it combines very poorly with open plan offices since you bother the people around you.
Non-native Chinese-Script language learner - I use touchpad stroke-input on my macbook 50% of the time (and touchscreen stroke input on phone 100% of the time).
However note perhaps the simplicity of the latin alphabet(a small set of separated characters) lent itself well to a simple implementation on early computers. which got them out of numerical processing earlier than if they had to speak chinese or even something like arabic where the written language has advanced to the point where the connected cursive form was the only correct form. the rendering of which would be tricky for early computers. note that english has a connected cursive form, the art of which has just about been destroyed by computers.
1. I don't read(or speak) chinese. but the characters are composed of sub characters and stroke order that may provide hints or insight into the nature of the character. Or they may not, I don't know, I don't read chinese.
I'm not a historian. But in the history, actually there was an attempt to "romanize Chinese".[1] And it failed. It's a complex issue, mostly a political one, and to say there is one single reason that it failed would be a gross simplification.
My anecdote is:
1. Chinese is a very "easy-to-read, hard-to-write" writting system. It's compact and information dense compared to alphabet based system.
2. There are too many homophones. In a daily conversation, you provide extra context with body language and your tone. But in a written context you have no such tools. All you have is "icons".
I mentioned that I didn't know and that was ignorant on the matter, and I wanted to know what people thought what the reasons are.
I have a multi-racial family, friends from across the globe (including china), and I am far from racist.
I think your comment is the most ridiculous and rude comment I have seen on this site
Is it read or read?
I doubt you can simply map the pronunciation of the chinese spoken languages to the latin alphabet.
And that's only the words, you also need grammar and different languages have tenses that don't even exist in other languages. Just compare Hebrew to English.
After that you need to train a billion people, you still need to conserve the knowledge of language or all the historical written texts are lost.
BTW were I come from it's at least 26 letters.
That's more logic between the symbols of tree, wood and forrest than between the same english words, and it even looks like some kind of tree.
Spanish has lots of those accent marks, so it's more than 26 letters, it's also accent marks.
Maybe better to compare it to languages that also use letters to represent words but aren't from the same family.
Maybe Arabic.
My guess is that there's a combination of network effects (i.e., 1.4 billion people already use it), cultural identity, and inertia.
If we used something else instead of english as the Lingua Franca, some say lobjan, maybe we could be more effective. Let's all switch.
I hope the point is clear. It's not all about the most efficient from your perspective but cultural, historical, pragmatic reasons, including that languages and ways of communicating mutate by themselves instead of being pushed top down.
It is so much more comprehensive - you never have to fumble about with the right way to pronounce something. Any particular combination of letters has one definitive way to pronounce it - so if you can read a phrase you can also speak it out without any ambiguity.
So much better than the mess english is. </sarcasm>
In particular, the phonetics of Mandarin Chinese underwent several waves of simplification to the point that many characters are pronounced pretty much the same - in particular, there are a lot of syllables pronounced /yi/ or /shi/.
So, transition to a purely alphabetic writing system would mean losing access to all the sophisticated texts of culture. There is even a poem illustrating that phenomenon, and taking it to the extreme: https://en.wikipedia.org/wiki/Lion-Eating_Poet_in_the_Stone_...
More practically, everyone learns their first language as a child, and at that point does not get to decide whether something is too "crazy" to learn or not, since nobody asks their opinion.
Further simplification was attempted at some point by the Communists (also as a means to increase adult literacy) but they rolled it back quickly.
Also, it's not 20,000 "icons" to learn. There are a couple of hundred composing elements ("radicals," although it's not entirely correct to call all of them this), which just repeat themselves in different arrangements, and there are some rules to it. Beyond these, only a hundred or so characters have purely unique elements.
One may think that is a disadvantage for Chinese, but since there such a wide variety of spoken Chinese languages, Chinese writing acts as a unifying framework across different spoken languages.
This allows a (more or less, ignoring traditional and simplified) common script for one billion people.
So a person in Taiwan can reasonably read a newspaper in Taiwan, Hong Kong, Shanghai, and Beijing. But that same person would probably have trouble speaking Cantonese with Hong Kongers.
There is a similar thing going on with written arabic, where Ḥarakāt diacritics indicate short vowels, long consonants, and some other vocalizations, but these are generally left out of writing except in the Qur'an. So you have a classic phonetic key with which to recite the Qur'an, but almost all written text besides that can be read by wildly different speakers.
I'm amazed at how two (Mainland) Chinese people can always communicate.
The level of shared cultural foundation is beyond what we have in Europe, I believe.
The Chinese language is an absolute joy to learn.
Also, several languages have moved from a script if iconographic characters to phonetic ones: for example Vietnamese and Korean, the former adopting a phonetic western alphabet with accents, and the latter developing a new phonetic script, hangul. Japanese developed two new syllables based scripts, hiragana and katakana, which are mixed with the old word characters (kanji). All these used to use Chinese styles characters prior to the switch.
No reason Chinese couldn't do the same. In fact, it did… pinyin is a formal phonetic alphabet for Mandarin Chinese that's based on the western alphabet with additional marks denoting the tones of the words. It's only hard I'm an educational setting, but you could use it anywhere!
The pinyin phonetic alphabet only works for mandarin, while a unified written script applies beyond mandarin.
Learning written Chinese not only connects you across varying spoken Chinese languages, but also connects you with the rich history of Classical Chinese text.
The formal pinyin system is great, and should be used along side with the actual Chinese characters. But there is no reason to replace the rich written Chinese characters, which connects across space and time, with a narrow and hollow substitute.
Despite the apparent difficulty maybe the people of China like their language nevertheless.
Languages vary widely in their expressivity.
Also, I think language (and the script) is closely related to the culture as well. The people of China might not want to let go of that.
And what about the literature and everything that's already written down in their native language? It's not just a question of translating everything. Trust me, things get lost in translation.
Also, if we're going to have everybody in the world work with the same script, there are arguably better candidates such as Arabic and Sanskrit.
In Sanskrit, for example, there's only about 50 letters (49?), it's impossible to mispronounce, the whole gender-neutralization situation becomes irrelevant (because well, things like car, village, fruit, pencil etc have genders (of which there are 3 btw — masculine, feminine, and neuter, but they aren't used in the way you would think)).
> why haven't china moved to a western (?) alphabet based writing?
By the way, here's how this appears to us non-westerners — okay here comes this Western hero who thinks the rest of the world is nonsense and should be replaced by the superior Western™ system.
I think maybe it's English that would benefit from being written in ideograms! The popularity of emoji is partly due to their compactness.
There is perhaps a benefit to having several mutually unintelligible languages all "sort of" compile back to the same written text, but that benefit is increasingly being eroded as the version of Mandarin spoken in mainland China becomes the lingua franca, not just inside China but also in parts of the diaspora.
If everyone is more-or-less able to understand spoken Mandarin, then it's no big leap to codify a phonetic written representation of that pronunciation, whether using Zhuyin or Pinyin or something else. It's a totally achievable goal, and we know it's achievable because Vietnam, Korea and Japan already did it. Claiming it's impossible is just Chinese exceptionalism.
The real question is not whether it's possible, or whether it would make the language easier to learn, it's whether Chinese-speaking people - and in particular the government with an authoritarian rule over the education of the overwhelming majority of Chinese-speaking people - want to do it. And the answer is they do not. And so it persists.
There are many reasons the characters have not been replaced by an alphabet. Ultimately, it comes down to the weight of tradition. Chinese has had basically the same writing system for 2000+ years. The characters are deeply embedded in Chinese culture, and the structure of the language itself is closely bound up with the characters. Only young kids use the alphabet, as a stepping stone to learning the characters, so proposing to use the alphabet comes across as if you want to dumb down the language. There are over a billion people who have learned how to read/write the characters, so the system has huge inertia and buy-in. As others have noted, this is like switching from imperial to metric, but a million times harder and emotionally fraught.
I wanted to link an article written by a linguist, but I couldn't find it. One of the hurdles would be tones: in order to express them, one would have to use plenty of diacritics and westerners would not be able to pronounce them correctly anyway.
I also grew up in Canada, have British parents, and lived in the USA for seven years, so I'm used to switching from "Traditional English" to "Simplified English" too ;-)
I mean, ya, sometimes Google Translate uses Mainland China words (e.g., I just typed Bicycle and it returned 自行車 instead of 腳踏車). But I just tried typing 土豆片 (potato chips) and translating it to "English" and it returned "potato chips" instead of the British "crisps" too. This is not a big deal at all.
As a matter of fact, a special "Cross-Strait" dictionary was developed to deal with all the language differences, with nearly 6,000 words and 30,000 phrases: https://taiwantoday.tw/news.php?unit=10&post=19596
It hardly seems like "no big deal" to those who should know best.
For translating a single word, a dictionary that lists all possible meanings in the target language is the right tool. A translator (human or algorithm) can only guess or give you the list of possible translations. Google Translate has come up with a suitable proxy solution by adding the "More Translations" pane.
It translated to left and correct.
What followed was the driver asking me “left or right?” And me saying “correct!” And he would turn left in confusion I would start shouting “correct correct” while getting more frustrated. He was like “great I’m doing good” while I was pointing the other way.
Good times I eventually had a chat with him and looked up “turn right” after driving around a block twice using left turns lol
Ignorant person here (sorry if I'm asking a dumb question), but given the larger population of China to Taiwan, why is that not the correct thing to do? Or is this "Traditional Chinese" a language spoken only in Taiwan?
When translating to Dutch, I wouldn't expect it to output Dutch as it is spoken in Belgium, Suriname, or Limburg either; I'd expect it to output Dutch as the main body of speakers speak it. They should get their own language designation instead, if the software wants to support translating into variants of the language
It's not a spoken language. It's only for written text.
> only in Taiwan?
Only in Hong Kong, Taiwan, and oversea Taiwanese/HK communities.
Taiwanese is another language... but let's use your simplified terms for now. Well, yes, pretty much this. It's not that different from writting "fish and chips" as "fish and fries" in British English.
Who’s mainly responsible for there to this day being no ISO standard for transliterating Cantonese, and its conspicuous absence from Google Translate (despite the whopping 86 million speakers—consider that Translate supports, say, Welsh, Basque, Dhivehi, and many others that are spoken by fewer than even one million), is probably obvious. That entity is much helped by the issue being opaque to most foreigners, who are sufficiently baffled by a completely different writing system to just not care and bundle all languages written in Chinese script into one basket.
There's no ISO standard because there's no actual standard. Even just within Hong Kong there are multiple systems in use.
I wouldn’t be surprised if there were attempts to introduce a standardised way of romanising Yue, but certain groups with many votes considered this national interest as they try to maintain authority across a very diverse set of people speaking different languages.
Because otherwise the default explanation for why there's no ISO standard for Cantonese romanization is the same as why there's no ISO standard for romanizing Burmese, or Malaysian Jawi etc. There's little demand internationally for such a standard.
And if the Hong Kong government's standard, or the Guangdong government's standard (this one presumably preferred on the Mainland) is officially blessed by ISO, would people currently using something else really switch just because of ISO?
When saying “little demand internationally”, could you clarify what international demand was there for a standard for romanizing Mandarin and why that is somehow not applicable to Yue?
Said standard getting ignored by most Burmese people in favor of ad-hoc romanizations, similar to Cantonese speakers in Hong Kong and Guangdong mostly ignoring various standards published by their respective governments.
> Jawi is a dialect of a language spoken by fewer than a thousand people total.
Malaysian Jawi https://en.wikipedia.org/wiki/Jawi_script is not a dialect, but an Arabic-derived writing system for Malay. While not all Malay speakers use it, it's certainly more than a few thousand.
> what international demand was there for a standard for romanizing Mandarin and why that is somehow not applicable to Yue?
Mandarin is an official language of the UN, whereas Yue isn't.
Standard being ignored by laypeople doesn’t make it any less useful in formal contexts sensitive to ambiguity.
Taiwan is its own country. India has a similar population to China. Why not just translate into Urdu or Tamil instead, by that logic?
I used the word China in other comments because I didn't know what to call it otherwise, hoping that in the context of Taiwan it is clear what I mean
對岸 /duì àn/ 'opposite [side of the Taiwan] Strait' is a neutral term coined specifically to avoid the controversy.
If talking about politics, another neutral way could be to just say 北京 /Běijīng/.
Under no circumstances is Taiwan every seriously referred to as "China," not even "Republic of China," the official name in the constitution, is really used anymore. The passports now prominently say "Taiwan," and polling indicates that "Chinese" identity in Taiwan is dying swiftly with the settler colonialist KMT.
The sibling comment may be referring to political discussions between politicians, especially when talking with PRC officials. I'm not sure, I've almost never heard those terms used except by really weird super-KMT taxi drivers.
The parent comment that "Taiwan calls itself China" is simply incorrect, and the government mentioned, in the 70s, was a KMT totalitarian government that's been essentially overthrown as of the 90s.
It's a pointed issue that you'll often find heated responses from because Taiwan, for basically the first time in its history as a globally participant nation, is finally getting to establish its own identity, separate from the Dutch, the Qing, the Japanese, and the ROC/KMT settler-colonialists. There's great fear that the CPC, having failed to win a culture war here with han chauvinism / han supremacy, will simply resort to violence to imperialise the nation.
So it's both incorrect and disingenuous to say "Taiwan calls itself China."
And so does Tsai Ing-wen:
> “We don’t have a need to declare ourselves an independent state,” Tsai told the BBC. “We are an independent country already and we call ourselves the Republic of China, Taiwan.”[0]
So it's not factually incorrect. Of course the issue is fraught and just saying "the country on Formosa is called the Republic of China" would be completely deceptive, but nobody is doing that.
Talk about Chinese politics is always heated but I don't appreciate being accused of being disingenuous.
[0] https://www.theguardian.com/world/2020/jan/15/tsai-ing-wen-s...
Furthermore, though Tsai Ing-Wen is the president of the RoC, as I said, there's a rising independence movement separate from the settler-colonial government of the RoC. Among these people, and the majority of Taiwanese, Taiwan is "Taiwan," and "RoC" is at best a formality, at worse an unchosen government underwritten by an unchangeable constitution.
The new passport illustrates the point decisively: https://upload.wikimedia.org/wikipedia/commons/6/66/Republic...
If you're unintentionally muddying the waters that's one thing, but I react strongly because I strongly oppose anybody that assists the PRC's cultural imperialism.
It is not. Traditional Chinese is not used in mainland china. It was superceded by simplified Chinese in areas controlled by the Chinese communist party.
This is a rather basic piece of knowledge when discussing Taiwan and Chinese relations.
They speak the same language in China and Taiwan: Mandarin Chinese. The question is whether to use more Taiwanese-sounding phrases when traditional characters are requested.
This seems purely political and obviously divisive: Taiwan (& HK) has its own history and China doesn't want it to.
However, it is worth noting that there was basically an iron curtain between China and Taiwan for a long time after Chiang Kai-shek went there. A lot of IT terminology for example developed separately in Taiwan and China, so a lot of technical terms are different. The following website lets you look up scientific terms and convert them between china mainland and taiwenese.
Traditional characters are still used in Taiwan and Hong Kong, while simplified characters are standard on the mainland and in Singapore. However, traditional characters are still used in some niche cases on the mainland (such as decorations, names of shops and restaurants).
Whether a translation to "Traditional characters" should default to Taiwanese-flavored Mandarin is an open question.
Traditional Chinese is used in Taiwan, Hong Kong, Macau, Singapore, and to one degree or another in various Chinese immigrant communities around the world. Taiwan is obviously the 800 pound gorilla here.
Taiwan has come to develop its own style of written Chinese -- think US vs British English but more nationalistic. So the question is: when Google Translate converts to traditional characters should it also be translating into Taiwan's written patois [not to be confused with the Taiwanese dialect]? After all, Google Translate does not differentiate between Britain and the US in English (it sadly uses British). And there are other groups which use traditional without this patois (Hong Kong, say, though likely not for long), albeit in fewer numbers.
The author does have one really good point, and that is that Google Translate is mistakenly describing traditional as zh-TW, that is "traditional as used in Taiwan". As long as it insists on doing that, it should also be translating into Taiwan's written patois. But what it should be doing instead is describing traditional as zh-Hant.
I'm a bit confused—can you think of some examples where GT would default to British English translations over American English?
Italian("Uso l'acensore") -> "I use the lift"
Italian("Provo l'acensore") -> I try the elevator"
Wrong way around? I just checked and it seems to use American (sadly).
They addressed this:
> NOTE: The Google Translate menu only says "Chinese (Traditional)". However, if you pick the option, you will see the language code reflected in the URL is zh-TW, which means "Traditional Chinese as being used in Taiwan". The alternative option for Google to fix this problem is to officially drop zh-TW support and switch to an appropriate language code instead, such as zh-Hant.
Sure, zh-TW is somewhat misleading. But they nor say that parameter is a ISO 639 or RFC 5646 conformed.
I wrote a speech in English, and had Google Translate it to Spanish. This was around 2010. One of the lines was "There is electricity in the air." The translation read, "No hay electricidad en el aire." I speak Spanish well, and upon proofreading, I freaked out that it would negate the sense of the whole sentence, but it can, and it did.
What makes you think Google "would" fix it?
Facebook has sided with Mainland China during 2019 protest in Hong Kong simply because there are more Chinese working inside Facebook than people from HK.
Youtube has sided with mainland China to censor words in Chinese comments simply because the censorship team employ, of course Chinese.
Just because Google doesn't operate in China, doesn't mean its influence are not there.
( Don't even get me started on Apple about its supply chains )
The URL contains a language code, but that’s just an internal implementation detail.
Imagine a world in which Google Translate has one entry called “Latin” for anything using Latin script (Welsh, Irish, you name it) and another entry “Simplified Latin” for English. Just like that, translation “from TC to TC” (e.g., Cantonese to Hakka) is very much a thing!
A Mandarin speaker in Taiwan has a different accent and slightly different vocabulary than Mandarin speaker in Beijing, but neither of them would be able to converse with a Yue speaker any more than a Russian speaker would be able to converse with a Polish speaker.
Also, you’re very mistaken if you think all versions of English are mutually intelligible.
Second, the fact that they also speak Mandarin (because it was a mandatory lesson at school or they were otherwise forced by circumstances to learn it) does not mean there should be no support in Translate for their native language, which incidentally dwarfs many languages that Translate does support in number of speakers.
If you think that’s viable logic, you should try applying it to other languages with predominantly multi-language speakers. Start with Irish (98% of people in Ireland speak English, after all) and move on to Catalan, see where that gets you.
Glad to see some debates on the vocabs listed there:
* Yes, it is very IT oriented.
* Yes, I have been in some academic environment. (but seriously, 文本 is rarely heard in my life!)
* No, it is not supposed to be an exhaustive list of all things wrong nor _the_ correct list.
At the end of day, it is a short list of what I'd like to see it changed and thus I tracked it. We may not all agree on the translations, but I think we can agree on 1/55 is a pretty sad score to have.
--
Lastly, feel free to fork and run your copy if you'd like a different set of words / translations to be tracked!
https://github.com/itszero/hasgooglefixedityet
Thanks for my 5 mins fame on HN!
Simplified and Traditional are different (but closely related) writing systems that can be used to express the exact same text. They are not different linguistic variants of Chinese.
It looks like regardless of which writing system you choose, Google still translates into the most widely used variant of Chinese: Standard Chinese, as spoken on the mainland. It just writes the result using the requested characters. It doesn't decide to use more characteristically Taiwanese phrases if you select traditional characters.
In order to address the complaint here, Google could allow one to select the writing system (Simplified vs. Traditional) independently from the regional variety (Mainland, Taiwan, etc.).
It's much more clear in Mandarin itself, such as 官話. Note that some might translate this as "Traditional Chinese," separate from the typical meaning, when trying to separate "simplified characters" from "traditional characters." In English it's all a mess.
Also, you can write guanhua or beifanghua or "Standard Chinese" in both Simplified and Traditional charactersets. They are almost 1:1. You also used to be able to write Japanese, Vietnamese, Korean, and other languages in Traditional Chinese characters.
That's why we should be careful to separate Guanhua ("Standard Chinese") from Taiwanese Mandarin and etc. imo using the word "Chinese" in the descriptor just sows further confusion and in fact furthers the political goals of the CPC to cast a political blanket over all things even remotely descriptive as "Chinese," be they language, culture, race, heritage, or nationality.
"Standard Chinese," "Mandarin Chinese," or however you want to call it (there are several terms in Chinese itself that refer to essentially the same thing) is the official language both on the mainland and in Taiwan. They have slightly different variants, but both standards grew out of the same historical movement to standardize Chinese. Both the Republic of China and the People's Republic of China wanted to create a national standard to enable universal communication in the country, and the PRC essentially adopted the standard the ROC had been working on.
This is an accident of history and a quirk of the English translation. Though the official language in English is described as "Standard Chinese," in Mandarin it's written as 國語, which just means "national language." The National Language described is Taiwanese Mandarin, 華語, which translates in English usually as Mandarin Language. By some measures it's very similar but it's different enough to deserve a different name, and the efforts of Taiwanese people to separate their culture from that of the PRC, and their ROC ancestors that came from the territory of the PRC, is a valid effort and should be acknowledged. Give it another 40 years or so and the languages will be quite distinct, as Taiwanese consciously incorporate more Hokkien, Hakka, and the various indigenous languages into their version of Mandarin to form a unique cultural identity.
> The Mainland is a well established concept when talking about China.
And people will happily, and inaccurately, describe the UK as "Britain" or "England," much to the annoyance of people who prefer to maintain their own unique cultural identity within the UK. In Taiwan this issue is even more critical as the sovereignty of the nation is under attack from many angles, including linguistically and culturally. There is no country on earth called The Mainland, and there is no country on earth called China, and CPC efforts to engage in cultural imperialism by casting a wide net around the concept of China and Chinese, as well as lay claims to historical Chinese imperial territories, and furthermore to imply that Taiwan is a PRC territory by calling the PRC the "mainland" (thus Taiwan as a province or colony), should be actively resisted.
It's not an accident of history. The Nationalists and the Communists had many similar influences (e.g., Sun Yat-Sen is revered by both). One of the ideas they shared in common was the idea of creating a national language. The PRC basically adopted the standardized form of Chinese developed by the ROC.
> Give it another 40 years or so and the languages will be quite distinct
I'm not so sure about this, given the amount of cultural interchange across the strait. Speaking in a more Taiwanese way is actually quite popular in mainland China, where it's considered cute.
> There is no country on earth called The Mainland
I'm not going to get into the entire nationalist dispute here. "The Mainland" marks out a well defined geographic area that's not quite synonymous with the PRC.
I'm not sure their models have a concept of writing systems. I noticed that neither Google Translate nor DeepL understand the equivalence of katakana and hiragana in Japanese. Probably every character is just an independent token to them.
Also, there is no reason for constantly calling it "Mainland China" where you could just as well call it "China," unless it is to further a political agenda.
Anyway, this is all beside the point. The concept of a "standard" language is political not linguistic. If Google is run as a business, not a political entity, their linguistic choices should reflect the language actually being used in any given market, and not be based on purported "standards" promulgated elsewhere. The same simple concept that somehow already works well for other language pairs that could be construed as similar to the point of being the same should also be applied here.
This is like saying the Canadians and the Americans each have their own unique language. They speak the same language, with small dialectal differences, and small differences in official standards (semi-official, in the case of the US). The internal differences in Mandarin as spoken in different regions of China are far larger than the differences between the ROC and PRC standards.
In the case of the PRC and ROC (now commonly known as "Taiwan"), the PRC standard is derived from the ROC standard, so emphasizing the differences is somewhat strange. They're very closely related to one another.
> Also, there is no reason for constantly calling it "Mainland China" where you could just as well call it "China," unless it is to further a political agenda.
"Mainland China" and "China" are not synonymous. Mainland China comprises the provinces of the mainland and Hainan. However you define China, at a minimum, it also contains Hong Kong and Macau, which are not part of "Mainland China." Hong Kong notably uses traditional characters, and it has slightly different standards than Taiwan.
> If Google is run as a business, not a political entity, their linguistic choices should reflect the language actually being used in any given market, and not be based on purported "standards" promulgated elsewhere.
Google's Chinese translations are so utterly terrible that this entire discussion is almost moot. I seriously doubt that Google is trying very hard to adhere to any particular standard version of Chinese. If you want decent Chinese <-> English translations, use DeepL.
I don't think the website is trying to "emphasize the differences," just point out the issues with Google Translate: namely that the output for "zh-tw" does not reflect the language actually used by people in Taiwan, and to that end it betrays the trust of the user. Of course it only lists where the problems are, so it's not a balanced view by definition. It focuses on what needs to be fixed.
In particular, as similar as the two languages or variants are to each other, nearly all the vocabulary relating to modern technologies developed separately, and is fairly distinct. I've never tried it but I can imagine a Google-translated text heavy on computer-related vocabulary can easily end up being unintelligible to a Taiwanese, which constitutes poor quality of service on Google's part.
Taiwan is a separate market for Google, and from the business perspective they would do best not to alienate their users there. Of course it is a free service with no reasonable expectation of quality. But if someone went to the trouble of listing all the issues, the problem might be worth addressing even for purely reputational reasons. I read through the whole word list and I'd say it's at least 95% accurate. Frankly, I'm surprised it sparked such a debate.
As for "Mainland China," you are technically correct about the scope. The term has its use in certain contexts if one is aiming to be very precise (or pedantic). But here it's tangential to the discussion.
There are dialect pairs where the question of whether they are separate languages is blurry. Taiwanese Mandarin and the official Mandarin of the PRC are nowhere near the level of difference where this question even arises. Mutual intelligibility is unproblematic.
> the output for "zh-tw" does not reflect the language actually used by people in Taiwan, and to that end it betrays the trust of the user.
If users were actually selecting "Taiwanese Mandarin," as opposed to "Chinese (Traditional)," and if Google were doing a decent job of translating into Chinese in the first place, I would agree with you. But neither is the case.
This on top of the fact that I'm pretty sure google maps uses some USA centric constant for traffic light timer estimations. Traffic light timers are very, very long in Taiwan compared to other countries. One time I had to wait 500 seconds (they had to add an extra big countdown timer to account for this). Usually it's more than 99 (the maximum the 2 digit counters can display). When you're in a car, motorcycle, or scooter, it's safe to multiply travel times by 1.5x to 2x. As for bicycle, I have no idea why google thinks bicycling is so fast in Taipei, but always multiply by at least 2x whatever time Google tells you.
Edit: the 500 seconds was an extreme outlier. Around 99 seconds is the usual in major Taipei intersections.
It doesn't make sense to delay thousands of cars on the highway by 30 seconds so that one car can cross it a few minutes faster. Better to make that one car wait 10 mins (and perhaps by now it is a queue of 3 cars), and save 30 seconds * thousands on the main road.
Other countries normally build bridges or highway merges to avoid this problem.
It's important especially when saying something is dumb, or telling someone they're foregoing basic due diligence, to spell correctly or the statement can lose it's poignancy.
languages don't have a strict bijective mapping between them
That, or a naive approach that attempts to map everything to mad-libs style sentences. See also; https://en.wikipedia.org/wiki/I_Can_Eat_Glass
[Un]fortunately, ja-en machine translation is dramatically better than it was when this website was launched, I guess it was more than 10 years ago. It used to devolve into really bizarre stuff.
This is a whole channel with exclusively Google Translate songs and creative content, folks.
Translation quality is one issue, e.g. "加油!" being translated to 'add oil!' instead of something like 'keep up the good work!' which is what it is supposed to mean in this context figuratively. But there is also a bigger cultural issue in that Google developers do not allow for people being multilingual: just because I made a choice for the interface language doesn't mean I need everything else translated to it (with no opt-out).
(and I just noticed that Deepl runs Linguee now)
And part of the reason come from these translation tools, there are many translated materials (website, documentations, etc) produced using tools that do not correctly handle the difference, and those content creators don't really care about that.
But this is also how culture works, we can and want to choose what to accept and what to reject, and just don't want to be the same as the ones who keep acting hostile to us.
Another thing is, related to the OP mainly focused on, they have tendency to oversimplifying words and merging words that are not related, which would cause problem when we communicate.
Take '質量' as mostly used example, which only means mass of matter in zh_TW, but it may also mean quantity and quality in zh_CN, and the list goes on and on.
And about those Japanese words, they don't pollute our tools like the ones in OP stated, we have to choose to use them.
edit: typo, more words
For example for "I want to build a good gym routine":
我想建立一個良好的健身計畫。 Wǒ xiǎng jiànlì yīgè liánghǎo de jiànshēn jìhuà.
Here are some variations of the sentence along with their English translations: 我想制定一套有效的健身計劃。 Wǒ xiǎng zhìdìng yī tào yǒuxiào de jiànshēn jìhuà. I want to create an effective workout plan.
我希望安排一個健康的健身日程。 Wǒ xīwàng ānpái yīgè jiànkāng de jiànshēn rìchéng. I hope to set up a healthy fitness schedule.
ps: I am building an app to make learning chinese easier. Feel free to ping me privately for testing :)
I find it fascinating that chatGPT is better at translating (a lot better) than a tool specifically built for it, when (if I understand correctly) chatGPT was in no way designed to translate, and is only in some way predicting text, one word at a time.
How come google hasn't leveraged better existing tech to make the translations better? Is it too computationally expensive?
I assume you’re right - that Google translate hasn’t been updated to take advantage of much bigger, more computationally complex models. I suppose for direct translation it’s probably not needed, but being able to ask chatgpt to explain the translation (and any cultural nuances involved) is a game changer when you’re trying to learn a language.
Which goes two ways: maybe this line of reasoning doesn't mean anything; or well yes exactly, but why so small online when they have all this space and also offer Bard.
(It's a bit confusing to talk about because surely it is just an older version of the same sort of thing, it's a less large language model right? I just think it could/should/would be a bit larger in the online hosted version.)
I assume the low latency you get from Google Translate is not feasible with current LLMs like ChatGPT. Translate is used to translate sentences on the go, live videos, translate entire web pages ; all of these would be too expensive (and slow) with an LLM... but things might change in the coming months/years as the tech improves.
I only ever use ChatGPT for translation now. GPT 4 especially blows everything else out of the water in quality.
I’ve done some insane round-trips through chains of totally unrelated languages and it retained 95% of the meaning. That’s superhuman quality.
Google Translate has been very broken for a while now and will often just completely refuse to translate, or only partially translate or translate in ways that are incoherent
Because you may have enabled that. I just tried this and it gave me an option to enable the feature to include surrounding text for better context or something.
As for translation, I get a tooltip to translate, web search or Wikipedia search
If ChatGPT outperforms google translate, I'd guess it's simply because chatGPT is a bigger model with more training data.
The engineers working on it seem to be focussing on supporting small languages (ie. languages that only a few million people speak - africa and india have lots of those, both big growth markets). They're also working on better apps - for example being to translate text inside images, text from a video feed, live translation of captions, etc.
It makes sense that they probably wouldn't be deploying a huge chunk of compute there to get a few marginal quality gains.
https://translate.google.com/?sl=auto&tl=sk&text=no&op=trans...
Anyway, the translation is wrong given the context as no numeral follows.
And it clearly isn't the translation anyone in the world would expect
Isn't Taiwan a part of China, though? I am not trying to justify Google not caring about Traditional Chinese as used in Taiwan. I am merely commenting as the author's view of China and Taiwan as two different entities while almost all countries don't recognize Taiwan as a separate entity.
And no, Taiwan is not a part of China (China says it is, but that's China...).
Given that the Chinese military doesn't control it, the answer it obviously no.