Unicode in five minutes (2013)
richardjharris.github.io
richardjharris.github.io
I do wonder if it is clear for people who are unfamiliar with Unicode? Anyone who is mostly unfamiliar with the details article covers who can say how comprehensible the article is?
I would also add a mention of the standard Unicode collation table that does a passable job for many languages at the same time (though Unicode Collation Algorithm is mentioned, which this is the default for, I think it's worth highlighting this property of most UCA implementations).
As for the article gotchas, multilingual text is even more complex when go past 5 minutes even for "simple" European scripts. Eg. in Bosnian/Croatian/Serbian in Roman/Latin alphabet, "nj" will be capitalized to "Nj" or "NJ" depending on the rest of the word — eg. "Njegoš" or "NJEGOŠ"; confusingly, Unicode also includes digraphs for both capitalization forms (the eternal tension in Unicode between encoding letters, glyphs or characters), even though they are linguistically equivalent — in practice, they are never used, which makes their inclusion even more perplexing (they are always spelled out using two characters, and there was no historical reason since none of the 8-bit encodings had them)! It will also sometimes be two distinct letters, especially in loanwords like "konjugovan" — this makes things harder when you need to collate texts since the proper order would be "konjugovan", "kontakt", "konj".
All of this is why I like to joke how Cyrillic script is technically much better for all of these languages, even though it is basically in official use only for the Serbian language — in Cyrillic, there is no conundrum in either of the above examples since nj=њ (or нј), Nj/NJ=Њ, and the order is clear: конјугован, контакт, коњ.
Slightly off topic, but just to riff on this a bit: maybe books and articles called "$THING in $NUMBER_OF $TIME_PERIODS" or "Learn $THING in $NUMBER_OF $TIME_PERIODS" should be retitled "$NUMBER_OF $TIME_PERIODS with $THING." It would be more accurate, not imply any sort of mastery, and, on top of that, sound a little more dignified. But, maybe it wouldn't sell as many books, so... ¯\_(ツ)_/¯.
https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
Disclaimer: might be biased because I've discovered Unicode through Spolsky's article.
Even if it's too topical to be actionable in every case, it gives you the general idea and vocabulary to put together useful search queries when you want to know more.
https://news.ycombinator.com/item?id=2035572
And the site:
http://unicodesnowmanforyou.com/
I wish I understood what the keepers of Unicode were thinking by including so much bloat in a character set (or character encoding). I realize that Unicode is going to have a huge number of symbols no matter what, if they're going to represent all the world's languages and math and punctuation, but I'd draw the line at emoticons, emojis, playing card symbols, and snowmen.
There was a lot of weird stuff in the world's character sets.
Emoji were first used by Japanese cell phone carriers. They were encoded as Shift JIS characters, but in incompatible ways. The Unicode Consortium had no real interest in this until Google and Apple basically said, "If we're going to have to support all these character sets, could we please standardize them?"
I think it's just the reality of standardizing the world's character sets. A lot of weird legacy stuff will slip in, and other countries will want to standardize things that seem unnecessary. Personally, I'm very thankful that somebody wants to do all the exhausting political work of coming to a consensus. A few snowmen are small price to pay.
And people use them as text, so there's a reason to add them and not much reason to refuse them.
What else would you question? Would it be more than 1500 more, which would bump it from 2 to 3 percent?
$ host xn--n3h.net
host: 'xn--n3h.net.' is not a legal IDNA2008 name (string contains a disallowed character), use +noidnout
Looks like emoji were forbidden in IDNA2008... :'(Magecart group uses homoglyph attacks to fool you into visiting malicious websites: https://www.zdnet.com/article/magecart-group-uses-homoglyph-...
Homoglyph attacks used in phishing campaign and Magecart attacks https://securityaffairs.co/wordpress/106916/hacking/homoglyp...
Forbidding mixed scripts fixes this attack. You also need to normalized names, and a few more minor things.
>>> [chr(0x07c0+i) for i in range(10)]
['߀', '߁', '߂', '߃', '߄', '߅', '߆', '߇', '߈', '߉']
0..9 in the N'Ko script BTW... py3> [chr(0x07c0+i) for i in range(10)]
['߀', '߁', '߂', '߃', '߄', '߅', '߆', '߇', '߈', '߉']
js> [...Array(10)].map((_,i)=>String.fromCodePoint(0x07c0+i))
['߀', '߁', '߂', '߃', '߄', '߅', '߆', '߇', '߈', '߉']I spoke way too soon. Unicode is weird. My apologies to our friend UncleEntity.
I must admit that I was surprised that the following snippet kept the LTR order in my terminal:
>> [(chr(ord('0')+i), chr(0x07c0+i)) for i in range(10)] [('0', '߀'), ('1', '߁'), ('2', '߂'), ('3', '߃'), ('4', '߄'), ('5', '߅'), ('6', '߆'), ('7', '߇'), ('8', '߈'), ('9', '߉')] >>> [(chr(0x07c0+i), chr(ord('0')+i)) for i in range(10)] [('߀', '0'), ('߁', '1'), ('߂', '2'), ('߃', '3'), ('߄', '4'), ('߅', '5'), ('߆', '6'), ('߇', '7'), ('߈', '8'), ('߉', '9')]
https://web.archive.org/web/20160417233039/http://babelstone...
But they did see fit to have ɑ (LATIN SMALL LETTER alpha)which is distinct from α (GREEK SMALL LETTER ALPHA).
unicode 88: http://www.unicode.org/history/Unicode88.pdf search for "A ligature is a glyph"
To start, consider that the term 'character' used in the article, though 'generally correct' ... is definitely not correct in the broadest sense.
Western, Cyrillic and Asian scripts boil down to 'characters' with some complexity maybe with ligatures ('Straße'), but it falls apart quickly for other languages.
Unfortunately, rather than creating rigorously applied definitions for things, and applying them consistently, even Unicode falls into this bureaucratic trap of vagaries with their own definitions.
So Unicode works well for most things, but then it falls off a cliff.
Here is the definitions section [1]
Even have a look at the definitions of 'Character' and 'Grapheme' and 'Grapheme Cluster' - and you start to see how confusion sets in very quickly.
Consider that in Unicode ... there isn't really such a thing as a 'character' - it's just an unspecific word we use that has no technical application! (When we say 'character' generally what we mean is 'Grapheme Cluster').
Language is itself a rabbit hole of complexity, so any standard trying to manage it will be painful - but it feels as though the true corner cases of Unicode are actually unbounded.
In short, too many pragmatic loose ends. Given any scenario where you think you have an alg sorted out ... and probably there are holes in it if you cared to try to find them for a specific language.
It's not 'bad', but it's not the uber solution, it's frayed at the edges.
> Consider that in Unicode ... there isn't really such a thing as a 'character'
This is a really important consideration, since it helps you realize the immense difficulty of wrapping your own logic for character-aware handling - unless you are deliberately limiting your scope, like only handling NFC-normalized text of a limited number of languages.