Unicode nearing 50% of the web
googleblog.blogspot.com
googleblog.blogspot.com
For example, some fine upstanding gentlemen decided that in Unicode (and GB-18030), Mongolian ᠣ and ᠤ, which are printed/handwritten exactly the same, shall be two "different letters" U+1823 and U+1824, but the different forms of ᠳ are the "same letter" U+1833. (And of course, there's ᡩ U+1869 which looks like what you want for some forms of U+1833, but you're not supposed to use it because it's "only for Xibe").
The closest analogy I can give in English is an encoding which forced you to use different "a" codepoints for the characters in apple vs. fake because of their different pronunciations, while making a single codepoint for "k" and "ck" and "c" (but only sometimes) because they sound the same. If you ever saw an encoding like that, you'd no doubt say to yourself: "WTF? I'm not using this, I'll stick with ASCII/EBCDIC/Morse code, thank you very much".
There are problems with unicode - ok, let's resolve them then. I still want to be able to address my email to the real name of person named in language A, living under address of country B, signing the email properly in language C. (where all parts use language-specific characters) Unicode is the first standard which allows me to do that in most cases, so I guess it's a step in the right direction.
As far as I know, Ge'ez (for Ethiopian languages) sets the record with 70+ encodings [1] which took all sorts of different approaches.
The Unicode design process worked out quite happily for Latin and Cyrillic alphabet users because
1. There was widespread agreement about what is the smallest indivisible unit of the script (thanks to long history of literacy education, decades of typewriter usage, etc.). No one suggested brilliant schemes like encoding "O" as "C" plus a right-concave combining mark ")" or I as "T" plus an underline, for example.
2. Among the hundreds of millions of users of those scripts, there were enough countries which had a reasonable history not just of typewriter usage but also of computer usage, enough time for them to develop various competing encodings whose mistakes Unicode could learn from
Inner Mongolian script pretty much presented the worst-case scenario compared to the above criteria:
1. The actual users of the script were a small and poor population with high illiteracy rates and not many computer users; and unlike e.g. Cambodians or Ethiopians, they had no big diaspora population of refugees living in the US or other high-tech countries either (hence no one fluent in English to advocate for them and point out problems in the proposed encodings).
2. As a result of #1, disproportionate amount of discussion surrounding the encoding was generated by scholars whose main aim was digitising quirky classical texts, not everyday people who wanted to write everyday things without the computer making them think of extraneous details they don't think of when they're writing by hand.
3. These scholars can't even agree what is the basic unit of the script (in Russian grad schools, they teach it as an alphabet; in Japanese grad schools, they teach it as a syllabary)
It's good news that Google are now decomposing ligature codepoints, although I do wish they had a version of their search that was literal; especially with programming-related and other technical searches, the special characters it filters out are often crucial.
The kind of massive indexes needed to support it must be insanely unprofitable given its tiny userbase in a tiny market.
It can represent all of Unicode while remaining backwards compatible with ASCII and C string representations.
The drawback, that characters have variable encoding, is almost irrelevant, because if you are using character indices, you are almost certainly doing it wrong anyway.
ALTER DATABASE database_name DEFAULT CHARACTER SET utf8 COLLATE utf8_general_ci;
And you don't have to worry about it for any new tables. (If you have existing ones, you'll have to change the table default, and possibly the column too.)
I generally agree.
Note: I think if you leave out any specific encoding configuration from your my.cnf, mysql implicitly goes with latin1.
For users who are assuming a modern unicode/utf8 default setup, might be nice if the mysql folks would require an explicit config for this setting.
Since we're not seeing an increase in UTF-16 or UCS-2 or anything else like that, this would seem evidence that the Web, or at least Google's view of it, is becoming even more increasingly dominated by Western European languages, which is itself an interesting idea.
This assumes that new pages (in languages with non-Latin scripts) are likely to use national encodings, which (in my experience) is not true. Given any popular* website written in Arabic, Cyrillic, Hangul, Kanji, etc, chances are that it will be in UTF-8. The only times I see encodings like SJIS, KOI8, or ISO-8859-1 any more are in old old pages, written before mass popularity of internet, or on amateur/very small websites. The former were created before UTF-8, the latter by unskilled or inexperienced authors. U+0000 to U+007F: 1 byte for UTF-8, 2 for UTF-16
U+0080 to U+07FF: 2 bytes for UTF-8 and UTF-16
U+0800 to U+FFFF: 3 bytes for UTF-8, 2 for UTF-16
U+10000 and up: 4 bytes for UTF-8 and UTF-16
Arabic, Cyrillic (Russian), Tamil, Thai, Hebrew, and many other non-European scripts take an equal amount of space in either UTF-8 and UTF-16. Japanese and Korean have some cases where UTF-16 is more efficient, but only by a byte, and most Chinese characters take 4 bytes either way.For network transmission, even gzip compression will take the sting out.
Oh, I know it's not free. I'm not saying it's free. I'm just saying it's nowhere near the pain point for most people. The days of using a computer and fearing that you might type a report for school so long the machine will run out of memory are long gone.
UTF-16 is a lot more work to support. For example, do I store the BOM in every string when I store a name in a database?
It's harder to handle UTF-16 because of all the nulls.
And since Chinese stores an entire word in a single character, I'm perfectly fine with their words taking 4 bytes. An average word in English uses 5 letters.