Why Unicode Won’t Work on the Internet (2001)
hastingsresearch.com
hastingsresearch.com
In short, Unicode will work just fine on the internet in 2016 as far as encoding all the characters goes. Problems having to do with how ordinal numbers are used, right-to-left languages, upper-case/lower-case anomalies, different glyphs being used for the same letter depending on the letter's position in the word (and many other realities of language and script differences) all need to be in the forefront of a developer's mind when trying to build a multi-lingual site.
In actuality we don't concern ourselves overmuch with planes, "octet blocks", and so on these days. There's just a single numbered list of characters* that we call "Unicode", and a few fairly minimal algorithms for transforming a list of characters into a bytestream ("UTF-8", "UTF-16", etc).
Still I love it as a historically interesting document. It's jarring to see "Oriental" used, I suspect the same author would never use that word in a professional context today.
[*] I use "characters" here when "code points" would be more accurate, but the details of that comparison aren't meaningful to this argument. I also don't go into normalization and so on, as this article seems to be more about the feasibility of fitting >100K code points into actual-bits-on-the-wire.
This expansion happened sort-of gradually. Unicode 3.1 (2001) assigned the Unicode 3.0 character set to be "plane 0", and added 14 additional planes‡‡‡ (for a total address space of log₂(15×2¹⁶)≈19.9bits). I'm not sure exactly which version between 3.1 and 9.0 the additional two planes in.
That is to say, Unicode 3.1 solved the address-space problem.
So why does the article have a section "Why Unicode 3.1 Does Not Solve the Problem"? Well, there are two answers: the one that the article suggests, and the one giving the author the benefit of the doubt.
The way the section is written, it seems that the author thinks the the address space of 16 bits + 16 bits is 2×2¹⁶ instead of 2^(2×16); because they are in separate "16 bit blocks". They seem to think that maybe 1 bit was added, but think that that just grew the encoding size by 16 bits. It honestly read to me like it was written by a linguist who only had a rudimentary grasp of programming; I was surprised when it said the author was a programmer at the bottom.
Giving him the benefit of the doubt that the argument was only poorly expressed, not poorly thought: Unicode 3.1 expanded the address space by another 917,504 encodable characters; more than enough! However, it didn't actually define 900,000 more characters; it only defined 44,946 more characters (as the author noted). But it wasn't limited to 16 bits anymore; it had a full 19 (more actually; 19.9-ish!) to work with. The author even mentions that 18 bits would have been plenty. Well Unicode 3.1 got them! They just weren't allocated yet.
That said, Unicode 9.0 (2016) still only has about 128,000 characters defined in it. A far cry from the author's claim of 170,000 characters needed to satisfy asian languages.
‡: To say otherwise would be to confuse Unicode with its encodings, a mistake that only leads to confusion.
‡‡: The term "plane" comes from ISO/IEC standards dealing with character sets. Unicode 3.0 corresponded to the ISO/IEC "Basic Multilingual Plane"; so each 16-bit group of characters got its own cutesy name as a "Plane" to match.
‡‡‡: Why 15 planes, then 17? It has to do with what was encodable with existing encodings. That isn't to say that Unicode was limited by the encodings; but that it was informed by them. The growth beyond plane 0 meant that UCS-2 had to be phased out for UTF-16 (its successor), as UCS-2 couldn't encode anything but plane 0. However, seeing that UTF-16 could only encode 17×2¹⁶ characters, it made sense to limit the number of planes to 17 if there isn't a pressing need for more; as doing so would require obsoleting UTF-16. And given that the current address space is only about 12% utilized, there's no reason to mandate phasing out UTF-16 yet.
As a side note, if anybody here is on OS X and wants to be able to type these characters, years ago I wrote a DefaultKeyBinding.dict file that adds bindings for these and a lot more (including the greek alphabet). You can see the file at https://gist.github.com/kballard/7584246fa5d5fcb684e996ff095..., and if you want to use it, just put this in ~/Library/KeyBindings with the name DefaultKeyBinding.dict. Any running apps may have to be restarted to notice this.
There's an extension for Google Docs for Latex BTW.
> Problems............all need to be in the forefront of a developer's mind when trying to build a multi-lingual site.
It will work. Just fine though? It sounds like way too much work!
Manipulating text, though, is inherently nightmarish. No format can prevent that.
> The current permutation of Unicode gives a theoretical maximum of approximately 65,000 characters
No, UTF-16 enables a maximum of 2,097,152 characters (2^21).
> Clearly, 32 bits (4 octets) would have been more than adequate if they were a contiguous block. Indeed, "18 bits wide" (262,144 variations) would be enough to address the world’s characters if a contiguous block.
UTF-16 provides 21 bits, 3 more than the author wants.
Except they're not “in a contiguous block”:
> But two separate 16 bit blocks do not solve the problem at all.
The author doesn't explain why having multiple blocks is a problem. This works just fine, and has enabled Unicode to accommodate the hundreds of thousands of extra characters the author said it ought to.
Though maybe there's a hint in this later comment:
> One can easily formulate new standards using 4 octet blocks (ad infinitum) – but piggybacking them on top of Unicode 3.1 simply exacerbates the complexity of font mapping, as Unicode 3.1 has increased the complexity of UCS-2.
They would have preferred if backwards-compatibility had been broken and everyone switched to a new format that's like UTF-32/UCS-4, but not called Unicode, I guess?
Hell UTF-8 was devised in mid-1992 and presented at USENIX in 1993.
> No, UTF-16 enables a maximum of 2,097,152 characters (2^21).
And until 2003 (and RFC 3629 which neutered it to match UTF-16) UTF-8 enabled 2,147,483,648 characters (2^31).
UTF-16 was only made into a standard in 2000[0]
The first non-BMP characters weren't introduced until 2001[1]
From a historical perspective, less than a month after Unicode 3.1 was released officially, it was not clear that Unicode (the standard, not the technical implementation of the encoding) would actually work.
[0] - https://tools.ietf.org/html/rfc2781 [1] - http://www.unicode.org/reports/tr27/tr27-4.html
It only got its own IETF RFC in 2000, but the surrogate pair mechanism was first specified in 1996's Unicode 2.0.
You should exclude the surrogates but include the non-characters when tallying the characters. According to Unicode's 2013 Corridendum 9, "Noncharacters in the Unicode Standard are intended for internal use and have no standard interpretation when exchanged outside the context of internal use. However, they are not illegal in interchange nor do they cause ill-formed Unicode text".
The number of code points available for use as characters must, by definition, exclude the noncharacters. So, to expand on the original comment, Unicode defines 1114112 code points, of which 1112064 can be used in interchange, and 1111998 can be defined as characters. UTF-16 can only represent the 1112064 that are valid for interchange, and the 66 noncharacters should generally be avoided (especially U+FFFE).
It's still widely misunderstood today.
* PHP has functions named "utf8_encode()" and "utf8_decode()", when they should have been called "latin1_to_utf8_transcode()" and "utf8_to_latin1_transcode()"
* MySQL for the longest time used latin1 as a default character set, then introduced an insufficient character set called "utf8" which only allows up to 3 bytes, not enough for all possible utf8 encoded codepoints, then introduced a proper implementation called "utf8mb4".
* mysql connectors and client libraries often default their "client character set" setting to latin1, causing "silent" transcodes against the "server character set" and table column character sets. Also, because their "latin1" charset is more or less a binary-safe encoding, it is very easy to get double latin1-to-utf8 transcoded data in the database, something that often goes by unnoticed as long as data is merely received-inserted-selected-output to a browser, until you start to work on substrings or case insensitive searches etc.
* In Java, there are tons of methods that work on the boundary between bytes and characters that allows not specifying an encoding, which then silenty falls back to an almost randomly set system encoding
* Many languages such as Java, JavaScript and the unicode variants of win32 were unfortunately designed at a time where unicode characters could fit into 16bits, with the devastating result that the data type "char" is too small to store a single unicode character. It also plays hell on substring indexing.
In short, the APIs are stacked against the beginning programmer and doesn't make it obvious that when you go from working with abstract "characters" to byte streams, there is ALWAYS an encoding involved.
Rust, and recent versions of Python 3 (but not early versions of Python 3, and definitely not 2…) pass this test.
I believe that all of JavaScript, Java, C#, C, C++ … all fail.
(Frankly, I'm not sure anything in that list even has built-in functionality in the standard library for doing code-point iteration. You have to more or less write it yourself. I think C# comes the closest, by having some Unicode utility functions that make the job easier, but still doesn't directly let you do it.)
¹Code units are almost always, in my experience, the wrong layer to work at. One might argue that code points are still too low level, but this is a basic litmus test (I don't disagree that code points are often wrong, it's mostly a matter of what can I actually get from a language).
> try to reverse a Unicode string.
A good example of where even code points don't suffice.
Edit: fixed link
(I use U+1F4A9 for testing all my non-BMP needs)
[1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
However, it exposes the encoding directly as a sequence of 16-bit ints. In other words, if you iterate over a string or index it, you're getting those, and not codepoints (i.e. it doesn't account for surrogate pairs).
Note that this only applies to iteration and indexing. All string functions do understand surrogates properly.
On the other hand, C# (or rather .NET) has a way to iterate over text elements in a string. This is one level higher than code points, in that it folds combining characters: https://msdn.microsoft.com/en-us/library/system.globalizatio...
Lisps with Unicode support seem to. This is a case where the Common Lisp standard's reluctance to mandate certain things paid off bigtime.
I don't know, but what you wrote sounds right. I don't really think of them as Lisps, although they have some Lisp-like features.
In the XML module, no less. I'll get round to moving those out of there eventually.
Unicode doesn't have to be complex; it is simply a single map of characters (such as "a" or "ﷺ") to a codepoint (such as "97" or "65018").
"Encodings": Encodings predate Unicode, and we had to deal with encoding before Unicode was ever invented (eg. "is this text in iso-8859-1, or is it koi8-r?"). Unicode has actually simplified the problem of encodings, for two reasons:
1) it's increasingly likely that any text sent over the internet is in the UTF-8 transformation of Unicode.
2) there is a mapping for any encoding into a Unicode transformation such as UTF-16 or UTF-8, therefore you only need three routines for text work: transform_encoding_to_unicode, process_text_as_unicode, transform_unicode_to_encoding.
"Planes": Planes are an aspect of Unicode that you may not need to concern yourself with, if you work in UTF-8 (as most Unix software does). However if you're in the world of Windows, Java, or Javascript (which use UTF-16 instead), then you may need to know that characters in planes other than 0 are encoded as 4 bytes instead of 2 bytes. Hopefully the tools available in your language of choice hide this implementation detail from you!
Misc: Beyond that it's true that there's an infinitely complex system of rabbit holes (think normalization, hyphenation, character ordering, right-to-left, ligature decomposition, and so on) that you can delve into when learning i18n, however Unicode wasn't the cause of them; they all existed before Unicode was invented! They're simply the messy but rewarding aspect of dealing with the written word.
a very useful site especially when having to explain what utf8 is to other devs when working in a windows shop.
Surely you're flattering yourself.
If all you see around you is wchar_t and LPCWSTR then that is what unicode means.
Any critique of unicode while not assuming UTF-8, which allows for more than 1 million code points) is a bit suspect in my opinion. The biggest point against UTF-8 might be that it takes more space than 'local' encodings for asian languages.
What's expensive to store are images and sound, from an ever increasing number of devices at an ever higher resolution. The production and storage of text barely registers in comparison.
> This happens for pure text[nb 2] but actual documents often contain enough spaces and line terminators, numbers (digits 0–9), and HTML or XML or wiki markup characters, that they are shorter in UTF-8. For example, both the Japanese UTF-8 and the Hindi Unicode articles on Wikipedia take more space in UTF-16 than in UTF-8.[nb 3]
https://opensignal.com/reports/2016/08/global-state-of-the-m...
Many "developing" countries never even deployed 2G and dial-up to any great extent. They were simply too poor to build large-scale telephone network infrastructure. When they did start getting connectivity in the late 1990s and 2000s, they were able to skip straight to the latest generation of technology.
This is a pattern we see all over the world — the wealth advantage of "developed" nations is offset by their historical investment in infrastructure that is no longer state-of-the-art.
For example, the London Underground has been in operation for over a hundred years. The newest bits are great, but the oldest parts are hamstrung by design decisions made in the Victorian era. Whereas, when China builds a new metro, it's able to build every part of it to modern standards, applying the accumulated knowledge from building those earlier metros.
> Not everyone has a gigabit internet connection--large portions of the world are still operating on 2G wireless, or even dialup.
Sure, the majority of people have moved over to 3G or better worldwide, but there are still many areas where 2G is more common. We just did a deployment in India[1] which still has more 2G coverage than 3G. Performance on 2G connections was a requirement from our Indian business partners.
It's also worth noting that your link contains an implicit bias: it's measuring connections, not people. People with slower connections sometimes simply won't connect at all if your site doesn't perform on their connection, so this is always going to skew toward faster connections. Your link is also pretty vague on the actual statistics--given their claim that the vast majority of countries have > 3G availability 75% of the time, 25% of the majority of countries could not have > 3G availability, and if the vast minority country is India, that's hundreds of millions of people.
Yes, the majority of the world is on 3G or better, but the minority can still contain millions and millions of people.
[1] http://www.sensorly.com/map/2G-3G/IN/India/Vodafone/gsm_4040...
https://en.wikipedia.org/wiki/UTF-8#Compared_to_UTF-16
Advantages
* Byte encodings and UTF-8 are represented by byte arrays in programs, and often nothing needs to be done to a function when converting from a byte encoding to UTF-8. UTF-16 is represented by 16-bit word arrays, and converting to UTF-16 while maintaining compatibility with existing ASCII-based programs (such as was done with Windows) requires every API and data structure that takes a string to be duplicated, one version accepting byte strings and another version accepting UTF-16.
Text encoded in UTF-8 will be smaller than the same text encoded in UTF-16 if there are more code points below U+0080 than in the range U+0800..U+FFFF. This is true for all modern European languages.
Most communication and storage was designed for a stream of bytes. A UTF-16 string must use a pair of bytes for each code unit:
* * The order of those two bytes becomes an issue and must be specified in the UTF-16 protocol, such as with a byte order mark.
* * If an odd number of bytes is missing from UTF-16, the whole rest of the string will be meaningless text. Any bytes missing from UTF-8 will still allow the text to be recovered accurately starting with the next character after the missing bytes.
Disadvantages
* Characters U+0800 through U+FFFF use three bytes in UTF-8, but only two in UTF-16. As a result, text in (for example) Chinese, Japanese or Hindi will take more space in UTF-8 if there are more of these characters than there are ASCII characters. This happens for pure text[nb 2] but actual documents often contain enough spaces and line terminators, numbers (digits 0–9), and HTML or XML or wiki markup characters, that they are shorter in UTF-8. For example, both the Japanese UTF-8 and the Hindi Unicode articles on Wikipedia take more space in UTF-16 than in UTF-8.[nb 3]
So I'll grant you a point but could match that against similar problems in UTF-16 (bad surrogate pairs, surrogate singletons, BOM bombs, and the same invalid code units).
UTF-8's very encoding quickly beats this out of anyone who tries, whereas it's easy to eek by in UTF-16. The real problem is that the APIs allow such tomfoolery. (Some have historical excuses, I will grant, but new languages are still made that allow indexing into code units without it being obvious that this is probably not what the coder wants.)
For the most part, the developer shouldn't really care about the internal encoding of the string, but the language/library should also not expose that to them.
Unicode standardizer: "Your software needs to support more than one language at a time. Please use Unicode."
Developer: "I don't want to. I've already got a codepage that's designed for the language I care about, and Unicode will make everything take up too much space."
Unicode standardizer: "Here's BOCU-1, an encoding of Unicode that compresses monolingual text into nearly the same amount of space as your favorite codepage."
Developer: "Uh, thanks, but that's weird."
Unicode standardizer: "Yeah, never mind. How about you try this new encoding called UTF-8?"
Developer: "Oh, this works really well and I guess it's small enough. I'll use it."
At the top level of abstraction is the abstract character aka user perceived character and grapheme clusters (a sequence of coded characters that should be kept together). Mapping from code-points to abstract characters is not total, injective, or surjective.
Unicode is not just standard for encoding. It's also standard for representation, and handling text.
The exact semantics of UTF-8 string in all cases is something that only few programmers are able to comprehend. I know for sure that I don't and I don't know anyone who does. Interchanging and storing UTF-8 strings and hoping for the best is the standard practice and it works well, but it shows the overreach that Unicode standard is.
That sounds really weird to me. Does that sound right to any native Japanese speakers here?
I'm not a native speaker, but if I were to make an equally strange metaphor as the author, katakana feels like writing in all capital letters.
But I guess there are some Japanese sources that are written all in katakana.
Kanji (Chinese-style characters) aren't phonetic on the other hand, they map to several different sounds, and several Kanjis may map to the same sound (not necessarily a single syllable).
Additionally, depending on context, hiragana might not have a phonological reading. See the object marker 'wo'. Katakana is usually phonological, except for the same usage of object marker.
In summary, someone who says "Hiragana can form pictures but Katakana can only form sounds" is either a philosopher who practices calligraphy or misinformed.
Technically yes, but they don't matter. The modern hiragana and katakana syllabaries have intentional symmetry. Anything written in one can be written in the other, and is (stylistic unorthodox choices of syllabary, old or limited computers where only one set is available, Japanese Braille which makes no such distinction, etc.)
> Additionally, depending on context, hiragana might not have a phonological reading.
This can also be true of katakana in some situations.
As for the neurological claim, I did a quick search and found an article which cited:
> Uno (as cited in Itani 2001) and Saito (as cited in Kosaka & Tsuzuki) say that kana is primarily processed phonologically in the reading process, whereas kanji is processed semantically.
http://www.staff.amu.edu.pl/~inveling/pdf/Dyszy-18.pdf
This is a few steps removed from the actual citation, but it seems to support your disagreement with the article.
This ignores the fact that a significant percentage of everyday Japanese life is now loanwords, so you can't really dismiss them as not being part of the language.