Dark corners of Unicode (2015)
eev.ee
eev.ee
Often when you want to say character you really mean "grapheme cluster"[1]. A Devanagri consonant cluster (of one or more consonants) with optional vowel (and misc diacritics) is a grapheme cluster. A latin letter with accent mark(s) is a grapheme cluster. A Hangul Jamo (precomposed or otherwise) is a grapheme cluster. A flag or multicultural family emoji is a grapheme cluster.
When talking about characters, the question arises if a Hangul Jamo made of decomposed characters is one character or three. Or if the combination of a [devanagri consonant + virama] "character" and [devanagri consonant + vowel] "character" is one character or two. This is specified in the case of grapheme clusters, and this usually maps to where the notion of a character is important -- text selection and offsets in editing, etc. "code point" does not universally map to any tangible concept -- it's a concept made up for the sake of specifying unicode. You only care about code points when dealing with UTF32 strings or when implementing operations on unicode text. "glyph" is also sometimes what we mean when we say "character", though that's more useful on the rendering end of things.
[1]: to be pedantic, "extended grapheme cluster", because unicode gave a rigorous definition of grapheme cluster and later decided to change it.
In Python, Unicode string is an immutable sequence of Unicode code points. It has nothing to do with UTF32 (Python uses flexible internal representation). To get "user-perceived characters" (approximated by eXtended graphemes clusters):
chars = regex.findall('\X', unicode_text)
width_in_terminal = wcwidth.wcswidth(unicode_text)
In practice, whatever is produced by the default method of iterating over a string is often called a character in that programming language (code point, or UTF16 code unit, or even a byte).https://pypi.python.org/pypi/regex https://pypi.python.org/pypi/wcwidth
$width-in-terminal = $text.chars;
$codepoints = $text.codes;
$bytes = $text.encode('utf8').bytes;
The .chars method should be the fastest, because Perl 6 internally uses strings of fully composed characters (normalized form grapheme). It's much better than having to do regex hacks like in Python.As far as I can tell, it turned out to be correct, though. Perl 6 has its own Inocode normalization variant called NFG (https://design.perl6.org/S15.html#NFG):
"Formally Perl 6 graphemes are defined exactly according to Unicode Grapheme Cluster Boundaries at level "extended" (in contrast to "tailored" or "legacy"), see Unicode Standard Annex #29 UNICODE TEXT SEGMENTATION 3 Grapheme Cluster Boundaries>. This is the same as the Perl 5 character class
\X Match Unicode "eXtended grapheme cluster"
With NFG, strings start by being run through the normal NFC process, compressing any given character sequences into precomposed characters.Any graphemes remaining without precomposed characters, such as ậ or नि, are given their own internal designation to refer to them, at least 32 bits in length, in such a way that they avoid clashing with any potential future changes to Unicode. The mapping between these internal designations and graphemes in this form is not guaranteed constant, even between strings in the same process."
From that, I guess Perl 6 extends its mapping between NFG code points (an extension of Unicode code points) and Unicode graphemes clusters whenever it encounters a grapheme cluster it hasn't seen before. Ignoring performance concerns (might not be bad, but I'm not sure about that), that seems a nice approach.
$width-in-terminal = $text.chars;
^- very likely wrong.wcswidth computes the cell-width (or column-count) of a Unicode string, which is unrelated to the count of graphemes, EGCs or code points. For example, Latin characters are one cell/column wide, while for example many CJK characters occupy two cells/columns, while they are still one EGC.
A typical application is printing CJK things to a terminal mask, progress display or similar.
Right, my point is that the concept of a code point is seldom useful unless doing storage stuff with utf32 or implementing unicode algorithms. Python may expose an API of code points but that doesn't mean that it's meaningful.
Performance arguments can be made as to why the API should use code points instead of grapheme clusters, so there are legitimate reasons for Python (and Rust, and many other languages) to do so. Sometimes you just need some comparable notion of length and "number of code points" is acceptable.
However, you should be careful when writing code that confers meaning to the concept of a code point. A lot of code does this (using code points when they mean glyphs or grapheme clusters).
> Right, my point is that the concept of a code point
> is seldom useful unless doing storage stuff with
> utf32 or implementing unicode algorithms.
In XML land, where strings are almost always UTF8, XPath offers a string-to-codepoints() function that returns a sequence of integers, and a corresponding codepoints-to-string(). These two have been invaluable to me on many occasions when doing string manipulation gymnastics.What kind of string manipulation gymnastics? I'd be wary of using codepoints for string manipulation for anything other than algorithms where you are explicitly asked to (e.g. algorithms that implement operations from the unicode spec)
For example, one I know quite well from having implemented it in Python: the HTML5 color parsing algorithm (the one that turns even incredible junk strings like "chucknorris" into color values) requires, in step 7 of the parsing process, replacing any code point higher than U+FFFF with the sequence '00' (that's two instances of U+0030 DIGIT ZERO).
And personally I think code points, as the basic atomic units of Unicode, do make sense as the things strings are made up of; I wish Python had better support for identifying graphemes without third-party libraries, but since Unicode encodings all map back to code points it makes sense to me that a Unicode string is a sequence of those rather than a sequence of some more-complex concept.
From my original comment:
> You only care about code points when dealing with UTF32 strings or when implementing operations on unicode text.
These operations fall in the latter. It's still pretty niche. If an algorithm is defined explicitly in terms of code points this makes sense. Stuff starts falling apart when people assign meaning to code points and use it as a placeholder for other concepts like glyph or columns of grapheme cluster.
On a related note, one of the best examples that reveal the confusing intricacies of Unicode was a previous HN thread about encoding Bengali:
https://news.ycombinator.com/item?id=9219162
Lots of intelligent comments from many contributors in that high quality thread.
For me, the takeaway was the truly understanding Unicode requires holding 5 different layers of concepts in your head simultaneously and the oft-cited Joel Software "Absolute Minimum You Must Know About Unicode" only covers 2 of them. (Which to be fair to Joel, his title emphasizes "minimum to know" and not "everything to know"). Also, at the highest level of abstraction, intelligent people can have philosophical disagreements on _what_ to encode in Unicode.
Because ASCII abstracts onto bytes well but text does not, layering an abstraction for 'text' that decomposes into characters and strings is where the problem lies. That's where, for example, Python 2 ran into trouble. Perl's set of compromises made its text abstraction useful, but creating them required creating them to be a principle focus of the language.
A stream of bytes is a stream of bytes, short of magic there's no way to generate a correct interpretation but for a predesignated protocol. And the only way to get a predesignated protocol for text is to make a deep study of human language and to choose to live with some compromises and not to live with others.
MySQL claims to support utf8, but in reality, it doesn't. You need utf8mb4 to support certain common Kanji characters.
This company had spent untold thousands (possibly millions) trying to convert gigantic databases (and I don't use the term gigantic loosely...) from utf8 to utf8mb4 because some of their Japan-based clients were using Kanji.
Sounds easy right? Wrong. utf8mb4 comes with some technical "gotchas" (google it) that had delayed the attempt to change to it by almost a year.
Anyway, I found this pretty amusing, and got a huge paycheck to explain to them just how screwed they were.
Oh, boy.
Turns out that trying to include astral plane code points in whichever version of Bugzilla that Mozilla uses causes comments to be silently truncated! Because MySQL.
I filed that one in 2010; it got deduped against a bug originally filed in 2007; it is now 2016, and the bug is RESOLVED FIXED, and Mozilla's bugzilla still has the same problem.
https://bugzilla.mozilla.org/show_bug.cgi?id=405011
At least the original Thunderbird bug has been fixed.
When I worked at Mozilla I was on the MDN (developer.mozilla.org) team, and we had this inexplicable bug: articles can be categorized with tags, and both articles and tags are localizable for all the languages MDN supports. So, for example, English reference articles on CSS properties were tagged "CSS Reference", while French reference articles on CSS properties were tagged "CSS Référence".
And... sometimes an English article's page would show it as having the French ("Référence") tag, and sometimes the French article's page would show it as having the English article's tag.
Turns out, MySQL's case-insensitive UTF-8 collation treated "e" and "é" as the same character. We didn't know about that, and hadn't noticed because the tagging library we used worked around it. Until one day a new version of it didn't, and tags from one language would start showing on another language's articles (if the words were the same, aside from diacritics/accents on certain characters). Which led to this:
https://github.com/mozilla/kuma/blob/00fc05b101658f863f58d7f...
That's a custom MySQL collation, which MDN defines and installs, to work around MySQL's default inability to tell "e" and "é" apart.
I know InnoDB limits index sizes to 767 bytes, meaning VARCHAR(255) using utf8 can have all 255 characters indexed, but VARCHAR(255) using utf8mb4 can only index 191 characters (floor(767/4) == 191).
After a quick Google search, that seems to be the most common gotcha. What other gotchas did you have in mind?
To be honest, I just don't remember. There was something about something that made something scary to the PM who was in charge of it all? That is about the best I can come up with.
I want to say the needed to index more than 191 chars, but that seems like a stupid thing to say. Who needs to index that many chars?
If I remember, I'll edit :)
edit: I guess I should say I was consulted to do some unrelated things, then helped them with some MySQL stuff that came up towards the end of the contract, then the utf8mb4 stuff came up, and I spent some time going through it with them. It was not the main focus of the contract, which is part of why I don't remember it very well. Just something that came up in the day to day...
This is quite possibly the greatest pun I've ever encountered
It's interesting to me that this happened far enough back that the only time we tend to consider 'th' to be letterlike is when we need to explain pronunciation — e.g. when reaching a kid to read
Of course then you'd have to teach people to use that character when writing instead of just typing "ij". And one day, there will be someone who, for stylistic reasons, needs tight control over when it's rendered as "ij" and when as "ÿ"...
Her is a link with some nice information: http://www.uazone.org/multiling/euroml/annex02.html
Based on precedent, I'd expect the Unicode Consortium to declare it the radical of a Chinese character.
(it does exist in Unicode, but as a ligature -- U+0132 for majuscule, U+0133 for minuscule -- and decomposes to the two code points for 'i' and 'j', with use of the ligature discouraged)
For example, the Chinese character for the word Biang[1] can be describe with:
⿺辶⿳穴⿲月⿱⿲幺言幺⿲長馬長刂心
[1] https://en.wikipedia.org/wiki/Biangbiang_noodles#Chinese_cha...
I'm glad I don't have to try to type that with 8 fingers and 2 thumbs. Writing Chinese with a keyboard sounds insane: http://www.slate.com/articles/news_and_politics/explainer/20...
(Except that biang is not encoded in Unicode yet so you can't type it anyway.)
What's there to think about? How else would you input a script with more than 50k characters?
> You type phonetically in one script, then select a character in another script that might be pronounced in a similar way.
Sure. Japanese works the same way, you input in kana or romaji, then select the suitable kanji (or kanji sequence).
Of course it only works when you have a regular phonology, that would be completely impossible for english since by and large orthography and pronunciation have no relation.
In fact English is becoming the same way: when you input "apple" and choose the [U+1F34E RED APPLE; stripped from input on HN] suggestion you've done exactly the same thing.
The Wikipedia page just uses images for the characters, I can't seem to find any actual example.
In the pencil-and-paper era, we had no problems ever with this.
It's only because we tried to 'simplify' things that we ran into problems :-)
As an additional aside, the OP just talked about sorting words within the same language.
What about sorting across languages? (For example, names?)
(I was the one who finally prodded the relevant Unicode cttee into fixing this bug. They did all the heavy lifting of writing proposals and steering the change through though: Thanks Ken et al!)
http://www.unicode.org/L2/L2014/14173-emoji-skin-tone.pdf
(I was fine with unrealistic, inhuman, Simpsons-style yellow...) I imagine fine gradations of locale-dependent zero-width gender identity modifiers will be added at some point. Unicode is a horror-show that will be producing bugs for decades to come. Every time you see a bug caused by "\r\n" vs. "\n", double-encoded HTML entities, or "smart" quotes, remember that Unicode is orders of magnitude more complex.