Dark corners of Unicode
eev.ee
eev.ee
A similar (albeit rather simpler and more limited) problem is calendrical calculus, where people who have little to no grasp of how to perform correct date operations do complex calendar applications and fail spectacularly in some edge cases.
Call me crazy, but if you are dealing with text, have some time set for research before you start your development.
Either, neither, both. Some are intrinsic to Unicode's purpose of encoding human text, others are accidents of Unicode history, yet others are design decisions which could have gone other ways (which may or may not have been better)
> Call me crazy, but if you are dealing with text, have some time set for research before you start your development.
The issues being most people will have a hard time justifying a year of linguistic and calligraphic study before the project gets to start (whether employed or independent) and most languages have "string"-manipulation facilities which are easy, obvious and wrong.
The simplest character encoding you could hope to work with is something like a single-line calculator or vending machine display - fixed-width, no line breaks, just the Arabic numerals and maybe a decimal point, some mathematical symbols, or a limited English alphabet to display "INSERT CASH". Any featureset above that produces issues. Just line breaks alone are responsible for all sorts of strange behaviors.
I think it's a bit magical that we've managed to do so much with text given the starting situation. At each step - from early telegram encodings through the proliferation of emoji - the implementations had to codify a thing that was previously left open, develop rules around its use, etc. We've made language more systematic than it ever was in history, for the benefit of machines to parse and process it.
It's funny though, how people are not even aware of the fact that different languages are different until they see something break with Unicode. Casing and collation rules have been language-specific before. It's just that before lots of software didn't even attempt to do it right. Again, a world I'd rather not have back again.
Let's restrict the calculator to whole numbers only.
ruby: https://github.com/twitter/twitter-cldr-rb
javascript: https://github.com/twitter/twitter-cldr-js
Human written language is pretty complicated. The Unicode standards (including the Common Locale Data Repository, the Unicode Collation Algorithm, normalization forms, associated standards and algorithms, etc) -- is a pretty damn amazing approach to dealing with it. It's not perfect, but it's amazing it's as well-designed and complete as it is. It's also not easy to implement solutions based on the unicode standards from scratch, cause it's complicated.
> JavaScript’s string type is backed by a sequence of unsigned 16-bit integers, so it can’t hold any codepoint higher than U+FFFF and instead splits them into surrogate pairs.
You just contradicted yourself. Surrogate pairs is exactly what allows UTF-16 to encode any codepoint.
Once you start talking about in-memory representation, you need to agree on an encoding. UTF-8, UTF-16 being the most common. wchar_t could be UTF-16 or UCS-2.
Here's the relevant section of the Unicode FAQ on the subject:
> UCS-2 does not describe a data format distinct from UTF-16, because both use exactly the same 16-bit code unit representations. However, UCS-2 does not interpret surrogate code points, and thus cannot be used to conformantly represent supplementary characters.
A correct UTF-16 implementation would interpret surrogate code point, validate that they're paired and prevent access to either surrogate via string operations.
When you come across an invalid sequence while decoding particular input (like "\ud83c") then you generally have three choices: throw an exception, skip the invalid part, or replace it with a replacement character. The default JavaScript behaviors is to be lenient. But if you need more control over the decoding behavior then you can use StringView or TextDecoder which is part of this spec: https://encoding.spec.whatwg.org/
Which is exactly why they are not and can not be UTF-16.
> The default JavaScript behaviors is to be lenient.
The javascript behaviour is to have UCS2 "strings".
JavaScript has byte strings, not character strings.
I'd a sequence of code units instead of code points. Sadly that holds true for many string implementations in programming languages, often for history, compatibility, or efficiency reasons (e.g. C#could have done it right, being designed after the UCS-2/UTF-16 split, but they didn't for various reasons). So you get code unit sequences with a few functions tacked on top that add code point support.
I don't understand why emoji are width 1 either.. really the EastAsianWidth.txt from the Unicode standard needs to match fixed with terminal emulators.
I've been dealing with all of this recently in JOE: http://sourceforge.net/p/joe-editor/mercurial/ci/default/tre...
In particular JOE now finally renders combining characters correctly. It now stores a string for each character cell which includes the start character and any following combining characters. If any of them change, JOE re-emits the entire sequence.
But which characters are combining characters? I expect \p{Mn} and \p{Me}, but U+1160 - U+11FF needs to be included as well but isn't. It's crazy that these are not counted as combining characters. Now I'm going to have to check how zero-width joiner is handled in terminal emulators. JOE is not changing the start character after a joiner into a combining character, ugh..
'compatibility' normalizations are mainly for comparison (including indexing/search/retrieval) and sorting, although there might be other uses. But indeed you should not expect a 'compatibility' normalization to render the same as the not-normalized input that produced it under normalization.
The 'canonical' normalization outputs ought to render the same as de-normalized input, but rendering systems don't always get it quite right.
For the web, the WWW consortium recommends a canonical normalization.
The Unicode documentation on normalization forms is actually pretty readable and straightforward, for being a somewhat confusing topic. http://unicode.org/reports/tr15/
> Normalization Forms KC and KD [compatibility normalizations] must not be blindly applied to arbitrary text. Because they erase many formatting distinctions, they will prevent round-trip conversion to and from many legacy character sets, and unless supplanted by formatting markup, they may remove distinctions that are important to the semantics of the text. It is best to think of these Normalization Forms as being like uppercase or lowercase mappings: useful in certain contexts for identifying core meanings, but also performing modifications to the text that may not always be appropriate.
The compatibility normalizations are pretty damn useful for indexing/search/retrieval though. Anyone storing non-ascii text in Solr or ElasticSearch (etc) probably wants to be familiar with them -- as a general rule of thumb, you probably want to do a compatibility normalization before indexing and again on query input.
Top-voted answer uses NFD, one below it uses NFKD: http://stackoverflow.com/questions/517923/what-is-the-best-w...
NFKD: http://www.perlmonks.org/?node_id=835238
NFD: http://www.perlmonks.org/?node_id=1105025
NFD: http://www.perlmonks.org/?node_id=485681
NFD: http://drillio.com/en/software/java/remove-accent-diacritic/
NFKD: https://gist.github.com/j4mie/557354
Two and a half of the first six results blindly apply NFKD to arbitrary text. All of them use normalization.
Sad state of affairs.
If you do want to do this, you should know that it only makes sense in your own locale, and you shouldn't be surprised that the methods are somewhat ad-hoc (I'm not saying you shouldn't do this: I've done it myself).
But it's true, as far as i know, that there's no unicode standard way to 'strip accents', which is unfortunate because we sometimes do need to do it. Even if 'strip accents' is locale dependent, and may have no sensible answer in some locales, I think there are sensible ways to do it in some locales (certainly in English, for Latin characters at least), and I wish there were a recognized best practice standard for doing it that could be implemented identically in various languages (maybe there is and I don't know it?).
There are unicode standard ways to compare/sort strings ignoring accents, in at least some locales, which might get you there if you reverse engineered them and took them further.
At any rate, at the end of the day, you can't simply talk about 'unicode normalization' without talking about the four different unicode normalization forms (canonical and compatibility; decomposed and composed) -- if you do, you are definitely getting something wrong.
And also, unicode normalization forms are definitely _not_ intended to 'strip accents', that is not what they are for, they aren't the solution to that, even if the compatibility normalizations do it in some cases.
I disagree with the notion of Symbola not being a pretty font. As I mentioned here¹, the glyphs Symbola has for the Mathematical Alphanumeric Symbols block are quite beautiful². (It may help that I’m using the non-hinted version on a HiDPI display though… still that implies it will look even better when printed on paper with an inkjet or laser printer since they still produce more DPI than the typical HiDPI monitor).
――――――
¹ — https://news.ycombinator.com/item?id=10198620
² — http://f.cl.ly/items/2h2p0r1F1h2E1y2o2y0c/Screen%20Shot%2020...
Well, I installed the Symbola font as he suggested but I'm still seeing lots of Unicode lego in the article.
I'm using Windows 7 and the latest version of Firefox, and I set the Symbola as the default font in Firefox and unchecked the box that says, "Allow pages to choose their own fonts, instead of my selections above".
What could I be doing wrong? I would assume that if the author recommends Symbola font, he's checked that Symbola has representations for all the symbols he's using.
Anyways, firstly, there’s no need to uncheck 'Allow pages to choose their own fonts, instead of my selections above'. You can go ahead and let the page specify whatever fonts it wants. If you don’t have the font(s) specified in the webpage’s stylesheet, or the webpage contains a character that the current font doesn’t have a glyph for, your OS/browser will substitute it for another font on your system that does have a glyph for that character (if it exists), so you can recheck that option. In fact, you don’t have to explicitly even pick Symbola to be used as a font for any type of text in Firefox at all, since your OS should use font substitution automatically if any of those fonts chosen there don’t have a glyph for a character on whatever webpage you’re on. In fact, to begin with, it’s impossible for any one font to contain all of Unicode right now, since even OpenType fonts can only contain a maximum of 65,536 glyphs, while Unicode has more than 120,000 assigned codepoints, so font substitution is absolutely necessary (so you can change the fonts in Firefox back to the defaults if you like).
Secondly, after you install the font, you may have to restart the computer or close Firefox before it actually picks up on the new font.
Thirdly, the version of Symbola that was linked to in the article is an old one. I’d recommend this² one instead (covers more codepoints).
――――――
¹ — https://support.microsoft.com/en-us/kb/2729094
² — https://web.archive.org/web/20150625020428/http://users.teil...
My response would be this: http://xkcd.com/1576/
Unicode is merely as complex that which it encodes: human language.