This is just my opinion as a native Chinese and English speaker, but I'm sure other bilingual English/Chinese speakers would tend to agree.
So the reading advantage, if any, is not necessarily incompatible with using a phonetic alphabet with a small number of symbols.
The lack of precision (or higher ambiguity) might also be seen as a disadvantage in the highly technical world of today, although I very much appreciate it in literary and philosophical works.
Also apart from the actual reading, which may be a tie, I'm not sure the complexity cost overall, e.g. teaching, fonts, IMEs, etc. are actually worth it. What would a designed writing system look like? I think a properly phonetic/syllable approach like Hangul or even the Kana might be correct.
You're thinking in too broad of terms. A system representing syllables as units would be a bad choice for english, which has extremely intricate syllabic structure (consider "strengths", which is CCCVCCC). This makes a Hangul-like system a stretch, though possible, and a Kana-like system, wherein every possible syllable has its own character, completely impossible. Most[1] potential english syllables are unused or used so rarely that no one could ever be expected to know their character. Douglas Adams named a fictional person "Slartibartfast" -- as far as I know, the syllable "slar" (rhymes with far) has no other existence in English. How would he have written it in an English syllabary?
Chinese has so few possible syllables that enumerating them is quite easy, but it doesn't use them all either. Kana work in Japanese because the only legal syllable structures are CV, V, and N.
> What would a designed writing system look like?
Well, all writing systems are designed; none are naturally occuring. It's hard to know what you mean by this, but:
- A syllabary works fine when the phonology of the language allows for it
- Spanish is a good, though not perfect, example of an alphabetic writing system that corresponds closely to pronunciation (a minor wart would be that I believe 'b' and 'v' are the same sound; c/z in Spanish Spanish and c/s in latin american Spanish have a similar problem)
- The Cherokee syllabary was created from nothingness within living memory
- The design goal of Esperanto was probably similar to what you're thinking of
Anyway, big picture, a given writing system is not equally suitable for all languages, so it doesn't make sense to ask which approach is "correct". They're more and less workable. Alphabets (in which the basic idea is one character per phoneme in the language) are about as simple as it gets, since the number of characters necessary is V + C rather than the O(V * C) needed for a syllabary in the languages with the simplest phonological structure.
[1] I have no numbers for this; it's a guess.
By designed, I mean "if modern scientists and engineers designed" versus more primitive people sort of making stuff up and folks adding on. With an explicit goal for efficiency; not aiming to necessarily "look nice".
When I look at Hangul, I see a rather simple underlying set of principles. It seems to logically build up. Compared to say, the Cherokee one, which, by looking at it for a minute, doesn't appear to have any structure. It seems like they're more-or-less random symbols. Maybe there is some deeper design there but it isn't apparent at a quick glance.
But is Hangul's choice of symbols the best representation for human minds for reading? And for writing or algorithmically dealing with characters?
It's the other way around; Hangul is also composed of more-or-less random symbols with no deeper design. Hangul is an alphabetic system much like Spanish, except that the letters are arranged into two-dimensional square syllables, and the squares then placed in a line, instead of the letters being arranged into a one-dimensional line directly. That's just an artistic choice.
I agree that Chinese is not the easiest or best language when all practical aspects are considered, and this is one of the reason simplified Chinese was invented, but the social hurdle to completely adopt a different language or change it even more dramatically would be far bigger than the nuisances that the language may have.
I don't know Chinese however however i found the lack of any similarity with European languages fascinating.
My increasing ability to scan lines hasn't slowed my increasing disability to write with a pen! Despite my occasional flirtation with wubi.
- Even in a single dialect, many words sound exactly the same, but mean different things. Using a phonetic system would harm comprehension. See http://en.m.wikipedia.org/wiki/Lion-Eating_Poet_in_the_Stone... for an excellent illustration of this point.
Increasing literacy by slightly modifying the writing system is why the PRC created Simplified Chinese in the first place. Now a more radical simplification of just using Hanyu Pinyin directly and not having a computer transliterate is an obvious next step. Hanyu Pinyin is not intelligible to speakers of non-Mandarin dialects, but this shouldn't be a showstopper for a government that is capable of this: http://en.m.wikipedia.org/wiki/Guangdong_National_Language_R....
Then the advent of computers and typing made stroke count mostly irrelevant.
While it was an interesting experiment, the fact that Taiwan still has a higher literacy rate than Mainland China, and one comparable to the other developed countries using the latin alphabet, shows that traditional Chinese isn't that much of a hindrance.
Also interesting: the linguist who developed Pinyin is apparently still alive at age 108 -- http://en.wikipedia.org/wiki/Zhou_Youguang
Chinese is in fact a group of spoken languages, including Mandarin, Cantonese, and Southern Min. These are completely different languages at least in tones and pronunciations. Mandarin speakers could hardly talk to Cantonese speakers.
But all these languages share a common character set, so that different dialect speakers can communicate in the written form even they couldn't understand each other orally.
Traditional Japanese (and similarly Traditional Korean and Traditional Vietnamese) uses a subset of Chinese characters which they call "Kanji". And "writing down the characters" (Bitan, 筆談) had been a very effective way of communication between Chinese, Japanese, Korean, and Vietnamese until the last century.
> But all these languages share a common character set, so that different dialect speakers can communicate in the written form even they couldn't understand each other orally.
Since everyone in China is taught mandarin Chinese they tend to write to each other in mandarin.
I'm not a native of Shanghai but I still get to understand written Shanghainese. Because even though I can't _read_ it, I can somehow _interpret_ it. It uses the same Chinese characters and I already knew what these characters mean.
> Since everyone in China is taught mandarin Chinese they tend to write to each other in mandarin.
There are still people native to Cantonese, Min Nan, or Hakka who don't speak Mandarin.
I admit that I only got as far as Mandarin I in school, but it seems to me that the existence of classifiers in Chinese languages contradicts that.
English raw text: 11KB
Chinese raw text: 8.8KB (UTF-8), 6KB (GB18030)
English compressed with xz: 3.8KB
Chinese compressed with xz: 3.7KB (UTF-8), 3.3KB (GB18030)
Yes, the information density does seem to be a bit higher in Chinese. Which, of course, is obtained at the cost of an encoding algorithm (writing system) that is more difficult to implement (learn), encode (write), and decode (read). It's the typical tradeoff between processing efficiency and memory consumption.
I wonder how much time it takes for the average American to type or handwrite the UDHR in English, versus the average Chinese to type or handwrite the same document in Chinese. After all, hard drives are cheap today. Human brain processing power is not.
I can't provide much in the way of authentic information as to your last question, but anecdotally I find it easier to finger-type Chinese into my phone than English. A couple factors are at work:
- I don't adjust my style for the phone, and it's next to worthless at predicting what I want to say (in English). This leads to a lot of thumb-tiring selection of single letters.
- I don't speak fluent Chinese, so most of what I say is pretty basic stuff, which the predictive input has an easy time with. This leads into the third point,
- Using a pinyin input method, you can call up entire phrases by just putting in the first letter of each syllable, e.g. xx for 谢谢 or bhys for 不好意思 or ng for 那个. This is basically the equivalent of your computer automatically expanding IANAL into "I'm not a lawyer" whenever you type it in, except, for everything. (This is also true for typing on a computer, but I'm comfortable enough on a keyboard that I don't have problems typing English.)
By the same reasoning, using 8-bit ASCII to examine the informational properties of English text is also absurd. English can be comfortably expressed with only 6 bits per character, after all.
Anyway, I did include a GB-encoded version in my calculations. And once you compress the text, the charset-related difference becomes much smaller anyway.
English can be comfortably expressed in 5 bits per character, using 26 letters and the five symbols [. ',"] . Capitalization and numeric digits are nice-to-haves. ;)
Not many Westerner seems to appreciate that Chinese Hanzi (and Japanese Kanji) are really words made up of simpler components arranged into squares, instead of written sequentially and separated by spaces. And the shape of the components in the square provides additional encoding information not available in sequentially written languages.
I really get screwed when I have to fill out a bilingual form and they want me to describe something in N characters (which it says in both English AND Chinese...doh!). This is just bias against us laowai!
Edit - I am not meaning this in its worst excesses, though I am meaning it pejoratively.
Just because there are 4 bases on DNA doesn't mean anything about what makes a good written language.
This is because they are not letters in a language, at least not as we commonly understand letters or language.
Those are just the labels we have given them as they were some of the closest metaphors to hand.
We could have called them notes and it would work just as well.
However that would have no bearing on the best ways to write music.