I really get screwed when I have to fill out a bilingual form and they want me to describe something in N characters (which it says in both English AND Chinese...doh!). This is just bias against us laowai!
English raw text: 11KB
Chinese raw text: 8.8KB (UTF-8), 6KB (GB18030)
English compressed with xz: 3.8KB
Chinese compressed with xz: 3.7KB (UTF-8), 3.3KB (GB18030)
Yes, the information density does seem to be a bit higher in Chinese. Which, of course, is obtained at the cost of an encoding algorithm (writing system) that is more difficult to implement (learn), encode (write), and decode (read). It's the typical tradeoff between processing efficiency and memory consumption.
I wonder how much time it takes for the average American to type or handwrite the UDHR in English, versus the average Chinese to type or handwrite the same document in Chinese. After all, hard drives are cheap today. Human brain processing power is not.
I can't provide much in the way of authentic information as to your last question, but anecdotally I find it easier to finger-type Chinese into my phone than English. A couple factors are at work:
- I don't adjust my style for the phone, and it's next to worthless at predicting what I want to say (in English). This leads to a lot of thumb-tiring selection of single letters.
- I don't speak fluent Chinese, so most of what I say is pretty basic stuff, which the predictive input has an easy time with. This leads into the third point,
- Using a pinyin input method, you can call up entire phrases by just putting in the first letter of each syllable, e.g. xx for 谢谢 or bhys for 不好意思 or ng for 那个. This is basically the equivalent of your computer automatically expanding IANAL into "I'm not a lawyer" whenever you type it in, except, for everything. (This is also true for typing on a computer, but I'm comfortable enough on a keyboard that I don't have problems typing English.)
Not many Westerner seems to appreciate that Chinese Hanzi (and Japanese Kanji) are really words made up of simpler components arranged into squares, instead of written sequentially and separated by spaces. And the shape of the components in the square provides additional encoding information not available in sequentially written languages.
By the same reasoning, using 8-bit ASCII to examine the informational properties of English text is also absurd. English can be comfortably expressed with only 6 bits per character, after all.
Anyway, I did include a GB-encoded version in my calculations. And once you compress the text, the charset-related difference becomes much smaller anyway.
English can be comfortably expressed in 5 bits per character, using 26 letters and the five symbols [. ',"] . Capitalization and numeric digits are nice-to-haves. ;)
I admit that I only got as far as Mandarin I in school, but it seems to me that the existence of classifiers in Chinese languages contradicts that.