The Forgotten History of Chinese Keyboards
spectrum.ieee.org
spectrum.ieee.org
Chinese characters are a type of pictographs that have some characteristics of QR codes. In fact, there is indeed a word retrieval method called four-corner number, which quickly maps Chinese character graphics to four numbers through some simple formulas, which is especially suitable for one-way encoding and retrieval. For example, the four-corner number of "龍" is coded as 0121, and the code of "兲" is 1080 (please refer to https://github.com/chai2010/im4corner).
In addition, Chinese characters are actually more important as hieroglyphic shapes. For example, we have a "凹语言" (Wa-lang https://github.com/wa-lang/wa/ ) designed for WebAssembly (WASM for short, WebAssembly => WASM => Wa), in which the Chinese characters "凹" and WASM The logo is very similar, and there was even a pronunciation of "wa" in the past.
After the popularization of computers, the function input method has been greatly improved, but there is still a lot of input resistance. For example, in programming, frequent switching between Chinese character names and English keywords brings a loss of input efficiency. As a programmer, I hope Chinese users can continue to pay attention to and improve these in the future.
https://patents.google.com/patent/US2613795A/
The idea of searching for characters via parts is similar to how Cangjie input method selects characters from radicals. I read somewhere that Cangjie input method was indeed inspired by Ming Kwai typewriter, but I can't find the citation for it.
One of Mao's better ideas
An alphabet can be adapted to basically any language. You just have to map the letters to the sounds, and you're pretty much done.
By contrast, the Chinese writing system is adapted very specifically to the properties of Chinese language. Every syllable in Chinese has a meaning (or set of meanings), so every character represents one meaning (or a few). English does not have that structure: words can have very arbitrary syllables that don't have any meaning on their own. Chinese characters encode a meaning plus a sound, which is often reflected in how they're composed (i.e., a character will often be composed of two simpler characters, one of which has the correct meaning and one of which has the correct sound). Chinese words do not change form: there's no conjugation, no plural form, etc. As a consequence, the writing system has no way to deal with things like conjugation.
I have no idea how one would even begin trying to adapt Chinese characters to write English. On the other hand, it's relatively easy to come up with a way to write Chinese in any alphabet.
But yeah, the whole notion is kinda silly. Most writing systems in the world are developed from very few originals. E.g. for most of Eurasia, the source is either Egyptian hieroglyphs or the Shang Oracle bone script.
Not only did it propel literacy rates to basically 100%, but it added a phonetic component to the language
> it added a phonetic component to the language
Fanqie has been a thing since the 2nd century. Zhuyin was invented in 1913.
And simplification's only "arguable merit" is that it saves a fortune in ink at the expense of losing its historical roots. But guess what? We mostly use computers now. So great job Mao, now we have two competing standards. (Nod to XKCD).
Unrelated but to those of us who started with 繁體字, simplified just looks ugly. (龙 vs 龍)
For example 竜 is a fairly common simplification of 龍 and imo not nearly as ugly
And about losing the historical roots, I guess if you're interested in it, the characters will always be there and accessible for you to study. I'd be interested how much the average Joe from Taiwan really remembers about random characters' roots, composition and meaning. I know much more people from the mainland, and among them are people who don't give shit, and those who can also write a lot of traditional characters and give lectures about the origin of meaning of some character and whatnot.
Also, since this is about computers after all, I've seen a study a while ago about from mainland where they tested how many mistakes people make writing less common characters. There was a bar chart that went down between 10 and 20ish, then went up a bit and started to go down again at around 30. It was speculated that people in school still have to write a lot by hand, and during/after college that stops and everything has been digital for a decade now so people just forget again, but folks old enough to have used pen and paper for a couple decades just had enough practice. I wonder if this effect would be more or less pronounced with traditional characters.
I prefer simplified for the aesthetics alone. Traditional is cringe and ugly in typed form.
Both massive wins
The problem also can also happen when translating from English, if you think about all the surnames that are occupations, or names like "bill" or "lily." Capitalization usually helps disambiguate, but there's title case and all caps and people who never capitalize anything...
The cases where simplification has removed those cues are rare enough that the extra complexity of traditional characters is really not worth it.
I've never heard anyone claim that simplified characters are more difficult to learn, and it just seems false to me.
Not sure what you mean by this. Do you mean that it's less convenient for people that don't speak / read Chinese? Why would that be a relevant metric?
You may be missing that character standards have changed over time and that different writing styles (草书,行书) are implicitly simplifications. You can think of latin or Russian cursive as a simplification of the printed letters.
In practice, the phonetic component has been mangled / evolved over time, so simplification doesn't make things more or less difficult for students (be it 5 year old native speakers or 50 year old non native speakers).
Because of these difficulties, there is a long tradition of anglicising names of settlements to meaningless collections of letters which when read by an English speaker approximate vaguely to the original Gaelic name. Unfortunately this is not a reversible process - you can't look at a modern anglicised name and guess what the Gaelic is, in general.
Now while Gaelic has a tiny population of native speakers, there are millions of people who know some "map Gaelic" - that is, we can look at a map with Gaelic place names, and understand the elements. It doesn't work for towns and villages, but generally in the north, no-one bothered to anglicise the names of natural features, just the settlements - and walking is the most popular outdoor recreation in the UK, so we learn this when we read maps.
When the first SNP government of Scotland came in, they introduced bi-lingual road signs, even in areas where Gaelic is no longer spoken. There was and is complaint over this, but I found that things became much clearer. I could look at a placename like Machrihanish, and see that it is Machaire Shanais. I still don't know what Shanais means, but Machaire is a type of landscape that I know, so I immediately know that this is low-lying and grassy, and fairly level. I can do this for thousands of place names without being able to reliably tell how to pronounce the words - similar to the way that the pronounciation of a word indicated by a Chinese character can vary widely with the part of China, so that the pronounciation becomes quite secondary to communication.
However, the Chinese language has evolved alongside the characters for about 3000 years, and it's very difficult to just separate the two. A huge amount of culture is bound up with the characters. Not only that, but the Romanized writing system is viewed as something that only little children use (as an aid to learn the characters). Once you've put in the effort to learn the characters (as about a billion people have), it's very difficult to accept their replacement by what is viewed as a script for children.
I feel like a proper comparison would not be number of characters, but a kind of pixel-budget, assuming both meet a certain reading speed and accuracy rate.
收天下兵, 聚之咸陽, 銷以為鍾鐻金人十二, 重各千石, 置廷宮中. 一法度衡石丈尺. 車同軌. 書同文字.
was translated into He collected the weapons of All-Under-Heaven in Xianyang, and cast them into twelve bronze figures of the type of bell stands, each 1000 dan [about 30 tons] in weight, and displayed them in the palace. He unified the law, weights and measurements, standardized the axle width of carriages, and standardized the writing system.For example,
"去" (pronounced "Qú") is "going to the". "超市" (prounced "Chao Shi") is "supermarket" "去超市" (pronounced "Qú Chao Shi") is "going to the supermarket".
3 syllables vs 7 syllables.
To me, it seems that instead of composing letters into words to convey meaning, they have more letters that are mini-words unto themselves.
> No matter how fast or slow, how simple or complex, each language gravitated toward an average rate of 39.15 bits per second, they report today in Science Advances.
-- https://www.science.org/content/article/human-speech-may-hav...
The size of a word does not correlate with it's concept - I still posit that some languages can transfer concepts faster than others, minus our baud rate.
Edit: Or, perhaps I am not as gifted an English speaker as my bias has presumed :| For example, I had to lookup "syntagmatic".
For example, the characters comprising your example text starts like:
collect (收) [from] [all] soldiers (兵) under the sky (天下), gather (聚) at(?) (之) Xianyang (咸陽), melt (銷) and (以) become(?) (為) bell-stand (鍾鐻) metal (金) person (人) twelve (十二) ...
As you can see, the English "translation" is more like an annotated translation. E.g., the original doesn't say who did it, or what he collected from soldiers: we just inferred "weapon" because what else could be melted into statues?
Similarly, "standardized the axle width of carriages" is just: cart (車) same (同) axle width (軌). We're supposed to infer "standardized" because we are talking about the Emperor's deeds.
There is a book, «Classical Chinese for everyone: a guide for absolute beginners» by Bryan W. van Norden that is easy to read and gives a gentle introduction into Classical Chinese.
The old grammar and vocabulary coupled with the Chinese style of writing metaphorically with an abundant application of allusions and with the same Chinese characters having multiple unrelated meanings, makes Ancient Chinese texts very terse and notoriously difficult to understand even for the educated Chinese people.
It’s certainly denser, though. And I agree about the front-loading of learning. It’s like learning vi. An absolute pain at first, then very comfortable.
The Prospects for Chinese Writing Reform (2006)
https://sino-platonic.org/complete/spp171_chinese_writing_re...
It is cited frequently.
How did that work out for Korea when they switched to Hangul?
They rely purely on context to distinguish {"apples", "apologies"}, {"mayor", "market"}, {"stomach", "ship", "pear", "double"}, {"acting", "delays", "smoke"} so on and so forth if what I'm scrolling is right. There's no tonal or character distinction. That surely isn't great.
Gradually more books and newspapers followed suit, because pretty much everybody found that writing everything in Korean letters actually make communication less ambiguous and easier to understand. If your phrase is ambiguous between whether someone's offering apples or apologies, then you just change the word or add additional context to make it clear which one is being offered. It's no different from how English speakers deal with bear/bear, tear/tear, arm/arm, ground/ground, and so on.
And back when it was first introduced, it certainly did wonders for literacy. Although it should be noted that original Hangul was more phonemic wrt its contemporary Korean, and the letter shapes were a bit simpler as well.
Perhaps it’s difficult to render in tiny Latin alphabet font, but if you have any Japanese or Chinese study under you, you could read and reproduce that nearly instantly on sight.
In our timeline I highly doubt whether it was the main reason why general purpose computers didn't happen first in China or Japan.
(via https://news.ycombinator.com/item?id=40548356, but no comments there)
How strange.
More on topic: Considering how inefficient Chinese characters are in general (but especially evident in computing) as one of the few languages where characters have no direct relation to phonetics, I wonder why there hasn’t been an effort to modernize it similar to Hiragana in Japan. Well, considering how Chinese is basically Kanji, why not just adopt Japanese?
We are not in the 90s anymore. UTF-8 has been around for 32 years now. If you’re working for a system that has no UTF-8 support, you have a much bigger problem to worry about.
> characters have no direct relation to phonetics
Most characters are phono-semantic where one part of the character is a phonetic hint and the other is a semantic hint.
> modernize it similar to Hiragana
Hiragana isn’t and wasn’t intended to replace kanji (unless you are from the fringe Kanamozikai). It serves a different grammatical purpose and is complementary to the other two. Kana is useful for an agglutinating language like Japanese, but not Chinese languages.
FWIW, the Japanese did develop a kana-based system for Taiwanese during the occupation, but it was an abomination.[1]
The phrase "a language is a dialect with an army" often appears in topic of Asian languages, and causing frictions between CJK non-speakers wondering about compatibilities between the three and speakers showing near vile dissents to those questions. While I understand both sides of these sentiments, the situation is not ideal for both sides.
IMO, it might be weird to refer to these languages as "Beijing Tokyo Seoul" languages, but doing so occasionally(just occasionally) could create more tangible feel as to why these three seem to exist side by side so utterly disconnected against each others.
nit: It's not accurate to say that the characters have no direct relation to phonetics. Thousands of them are semanto-phonetic compounds, meaning they combine a character relating to the word's (or syllable's) meaning with a character relating to pronunciation. Sinitic languages tend to have a lot of homophones or near-homophones, so this approach works reasonably well as a memory aid once you've memorized a bunch of the basic characters.
One problem is that many of the pronunciations have drifted from the Middle Chinese pronunciation of the words. Also, some of them have been simplified in Simplified Chinese which makes the components a bit harder to discern.
I've been learning some Cantonese recently and this is very apparent with certain common Cantonese words. For example, the first-person pronoun in Cantonese is pronounced ngo, with a low-rising tone, and written like this:
我 https://www.cantonese.sheik.co.uk/dictionary/characters/1/
The word for goose in Cantonese is also "ngo", but with a different tone. Here's the character for that:
鵝 https://www.cantonese.sheik.co.uk/dictionary/characters/1200...
If you enlarge it, you'll see that the left side is the same 我 from before. The right side is 鳥, which means "bird" (https://www.cantonese.sheik.co.uk/dictionary/characters/161/). So if you saw this character and knew the basic characters for the pronouns and the word "bird", and you spoke Cantonese, you'd be able to easily understand what it meant.
Here's another one. The word "ngo" with still a different tone means "hungry". How do we write it?
餓: https://www.cantonese.sheik.co.uk/dictionary/characters/740/
In this one the phonetic component is on the right instead, which is a bit inconsistent. The left side is this:
食: https://www.cantonese.sheik.co.uk/dictionary/characters/116/
What does 食 mean? It's the verb "to eat". So if you saw this 餓 character and knew a couple of other basic characters, you could figure out that it's the word "ngo6" meaning "hungry". Many of the characters still work like this although the sound shift I mentioned above means that some work in some Chinese languages and not others.
I am working with other volunteers to improve Cantonese teaching, and wonder what difficulties you have encountered when learning Cantonese, and what materials or communities would be helpful for Cantonese learners.
Because Japanese characters have no direct relation to Chinese phonetics. Both belong to different dialect continuums, phonetics aren't compatible.
And I suspect same might explain lack of native Chinese phonetic script; `Chinese` isn't a single spoken language, but what is called as such is its Beijing area version of one of Chinese(or Sinitic) languages. The written language was universally understood in China due to bureaucratic needs, but AIUI it's not same as spoken language and it's not necessarily used everywhere. Maybe they just had little uses for a standardized phonetic script?
1: https://en.wikipedia.org/wiki/List_of_varieties_of_Chinese
1. That Chinese writing is inherently inefficient. It's actually very efficient...to read. And nothing beats the efficiency of having a script that maps perfectly to the language. Also as sibling comment notes, UTF-8 is a thing.
2. That there is no relation between written characters and phonetics. Incorrect, as several sibling comments point out.
3. That Japanese kana represents a successful "modernization" of kanji that Chinese should emulate.
4. That Chinese is "basically kanji" - assuming the Chinese and Japanese languages are essentially interchangeable. They...are not. I can't even begin to emphasize how much they are not. Chinese is subject-verb-object while Japanese is subject-object-verb, for instance. Chinese also has many phonemes that are incompatible with Japanese, which would not be covered in hiragana. Finally, kanji came from Chinese and has subtle differences and while it is mostly a subset of Chinese hanz, it has its own slightly different character set
Also it's not uncommon for words like ろ過(濾過)to be written in part kanji especially in news... if that trend continues beyond the 常用 kanji we might end up with a Japanese that is closer to Korean.
The Japanese economy has been stagnant for over 30 years with no end in sight. Following the same logic, Japan should perhaps “modernize” their language by following China, which is a ridiculous conclusion as you can tell.
Tokyo from Beijing(2000km/1200mi) is about as far out as Paris to Kyiv. Far East countries are also separated by seas, like Mediterranean countries across the sea. I doubt a lot of Parisians have meaningful ideas of "basically Latin" Ukrainian any way or form, or Italians with Tunisian, but there's such false instinct that forms out of above-mentioned presentation that those Asians are rather next door neighbors.
That and mistaking personal difficulties and inefficiencies associated with understanding languages in non-native manners as inferiority of the foreign one.
It's a failure to recognize that languages (which I would rank music a kind of) evolve organically, and outside of some edge cases, like Esperanto, they're not artificially created in a vacuum.