Unicode is hard
shkspr.mobi
shkspr.mobi
> ⨈Һ𝘢ʈ ╤ћᘓ 𝔽ᵁʗꗪ
Assuming the printer uses ESC/POS[0] (which is likely), the codepage is part of the printer's state. To change the code page, the driver sends a specific ESC command (<ESC t x> aka <1B 74 XX> where x/XX is the desired codepage byte) (none of which is "UTF8" incidentally) and you can change the codepage before each actually displayed character.
So it's the driver software fucking up and either misencoding its content (most likely) or selecting the wrong codepage. The £ might be displayed correctly on the right side because it's e.g. hard-coded (properly encoded) while the product label is dynamic and when that was added/changed no care was taken with respect to properly transcoding. The printer absolutely doesn't care, it just maps a byte to a glyph according to the currently selected codepage.
[0] ESC because the protocol is based on proprietary ESCape codes[1], POS because the entire thing's a giant piece of shit
[1] https://en.wikipedia.org/wiki/Escape_character#ASCII_escape_...
Why not? Shouldn't all printers made nowadays support UTF8? Probably by default?
Mind boggling that this is still a problem today.
Even in the travel world it goes wrong all the time. You'd expect large international travel organisations (yes, talking to you Tui!) to be able to handle UTF8 names since many of their customers and locations will have special characters, but no. I once was nearly refused to board an airplane because the name on my ticket did not match the one on my passport...
Sure, there are recommended practices but there have been enough mistakes already (or lazy programmers) that it is hard to be confident that any string with “interesting” symbols in it is exactly what it appears to be. And there have been security problems related to the fact that many interfaces expect the user to know exactly what they’re reading, much less the programmer.
Like how "е" looks identical to "e", which looks close to "ė" if you aren't careful, which might be mistaken for "é" in smaller fonts even though all 4 letters are different unicode glyphs.
Being able to say "ɡrеɡ" is the same as "greg" even though 3 of the 4 characters are actually different would be extremely useful in some cases, and in others would be extremely incorrect, so giving the developer the ability to say how "exact" they need their checks to be in a "native" and easy way might go a long way toward not only making this problem more "obvious" but also toward forcing them to be explicit about what they are checking for.
Do you group c with ç and č? In English you would. In France, Portugal, Serbia, the Baltic States or Czech republic you may not.
The ė mentioned is distinct from e in Lithuanian.
Even in those languages you still might want to treat an OCR of a "c" with "ç" in some cases, or you might want to treat them identical when moving from a system that only accepted ASCII in the past to one that is fully unicode compliant. And even in English there are situations where "ç" should not be treated as "c" (like in URL resolving or other scenarios where an exact match is needed).
It's not technically correct, but it's "more correct". And giving the developer the ability to determine where along that "scale" of exact-ness they want to be might help.
For suggest functions, it turns out people fully expect to be able to write some sort of ascii normalisation and still get a match, i.e. the address has ż in it, but people want to type a plain z.
And the rules for this are not entirely obvious. A swede would totally expect being able to write ö instead of the norwegian ø when doing routing across the border.
If you were granting e.g. domain names, or usernames, you'd be able to map each character in the test string to its homoglyph equivalence-class, and then ask whether anyone has previously registered a name using that sequence of equivalence-class values. So someone's registration of "тhe" would preclude registering "the", and vice-versa; but when you normalized "тhe", you'd still get "mhe".
Of course, to use such a system properly, you'd have to keep the original registered variant of the name around and use it URL slugs and the like (even if that means resorting to punycode), rather than trying to "canonicalize" the person's provided name through a normalization algorithm. Because they have "[the equivalence class of т]he", not "mhe"; someone else has "mhe".
I believe gp is talking about the font. In some fonts (especially italic/cursive), the letter "т" looks like "m", and nothing like "T" -- so it's really hard to say with which one it's "visually equivalent".
In the search space, therefore, when you index the word 'café', you also index 'cafe' with a smaller weight. And when you see the query [café], you expand the query to ('café' OR 'cafe'-with-smaller-weight)
And you don't want to do either of this if the two words are actually different!
As an example of this in the wild, the ElasticSearch docs talk about the issue: https://www.elastic.co/guide/en/elasticsearch/guide/current/...
PRECIS appears to be aimed more at figuring out if 2 usernames are 'the same'.
It would be nice if the authorities handling the registration of the domain names could forbid domains that look to much like each other.
However, there are ways around this too. I think the fundamental mistake was to allow (all?) unicode strings as urls. However, I can't come up with an elegant solution on the spot (since it would be unfair and unpractical to use ASCII for this).
Data file mentioned in the standard: http://www.unicode.org/Public/security/9.0.0/confusables.txt
Unicode specifications are incredibly thorough and well thought-out. The problem is that the Unicode spec isn't shippable software. It's not an implementation.
And there's no singular implementation. Worse, nobody uses any particular implementation the same way, and rarely to its fullest extent. Compounding the problems, so much code is _proprietary_. You have no way to verify and track how such code will behave, so interoperability is difficult. For example, good luck trying to reproduce the behavior of Outlook, Mail.app, and gmail.com in terms of how each will highlight URLs in free-form text.
The only saving grace appears to be that the rest of the world, I assume, has grown accustomed to how broken American software is in terms of dealing with I18N issues. And Americans remain blissfully naive. I keep waiting for the other shoe to drop; when managers will finally crack the whip at the behest of international customers and demand that engineers begin taking I18N seriously. But it hasn't happened yet. I've been waiting almost 15 years, accumulating skills and best practices that my employers don't seem to value very much. Oh well....
We're going to be suffering for that mistake for a looong time.
I can see why you'd say that, but who decides whether it should?
a+b=d+d
Shouldn't be uppercased.
It's an insoluble problem to put contextual semantic info into Unicode characters, because individual characters have no context.
It's used in typesetting sometimes, and if a character is used then it should have an encoding.
IMO there's little semantic difference so it doesn't deserve a character. We should have drawn the line between content and formatting, but it's too late and what we have now is emoji and one-use glyphs. [1]
No. In Japanese, how you read/pronounce a character depends on context. Sometimes they are the same as Chinese, sometimes not.
Take mountain (山) for example.
Using the Chinese pronouncation it is "san". 富士山 (Mount Fuji) is ふじさん "Fuji san"
Using Japanese pronouncation it is "yama". 山登り (Mountain Climbing) is やまのぼり "yamanoboru"
(and don't call me Shirley)
PS: Isn't り pronounced "ri" and る pronounced "ru"?
Yes, I typoed that and it's too late to fix it. り is "ri", not "ru".
exactly!
Isn't this why NFKC normalization exists?
You, as the programmer, need to understand each of them and why you want to use them.
Brief overview: https://en.wikipedia.org/wiki/Unicode_equivalence#Normalizat... More technical details: http://www.unicode.org/reports/tr15/
I --THINK-- offhand, that NFKC is what you want to use when preparing a password input for processing/comparison (it's lossless, but to a specific point). I also --THINK-- that NFC is the form you want to use when retaining source glyph language distinctions.
From the stackoverflow hits:
https://stackoverflow.com/questions/16173328/what-unicode-no...
I agree with the destructive (pre computation/comparison) operation and that either of the NFKD or NFKC forms should be used (since they destroy non-printing differences for visually compatible characters; a more user friendly approach).
The 'C' forms are always more condensed (accents are packed in to a single character where possible), and thus of higher entropy per input byte. It is my belief that this form is likely to be less susceptible to attacks.
The 'D' forms seem like good choices for /editors/ where the precise nature of a character might be altered by adding or removing accents. (Most human input boxes; during the input/edit process)
What sort of attacks are you talking about?
> The 'D' forms seem like good choices for /editors/ where the precise nature of a character might be altered by adding or removing accents. (Most human input boxes; during the input/edit process)
Editors shouldn't care about C vs D. The reason being, once you've typed the grapheme cluster, it's supposed to act in an editor as if it's a single "character" regardless of whether it's made from one codepoint or several. This means that if I type é then arrow keys and the delete key will operate it on it exactly the same whether it's composed or decomposed.
If the address gets somehow mangled, as it often does, the shipping label could read "Ro?u street" and the postal offices guess on the address being "Roņu" or "Rožu" is as good as anyones.
I too usually just write "Ronu street"
Because even with no characters messed up, many cities have 2 or more streets with the same name.
Coincidentally, at least here, the official addresses are guaranteed to be unique, a city/town can not have two streets with the same name (at least at any single point of time, renaming over centuries gets weird) - naming streets and assigning numbers to houses is not an arbitrary decision that the landowner can make, it is managed/coordinated centrally.
Also I'm much better remembering addresses of people instead of their GPS coordinates, especially when they're close together, like in a city, where the coordinates would all be almost the same.
And if they do write the wrong city, well, that's what the zip code is for.
I totalled the numbers up again and the total was exactly £1.50 less than the total shown on the bill! My wife (having no faith in my basic adding skills) pulled out her phone telling me "don't be silly" and added them up to get the same result as I had.
We asked the waiter about it, who disappeared off to get his own calculator.. He added things up, looked confused and then took it off to the manager. She then repeated the process on the calculator and also looked confused, unable to explain what had happened. They gave us £1.50 in cash, apologised and then kept the receipt (I guess they didn't want us posting that on twitter!).
To this day I've no idea what happened. You could suggest that some programmer somewhere is getting rich off this, but it seems rather unlikely to me. I'd really love to know what the cause was (and whether the manager ever reported it further up the chain; because this seems like a rather serious error to me.. how often does it happen? is it always £1.50? did the issue get found/fixed?).
If I remember correctly, most booze is at the full 20% rate. I can imagine a mistake like this happening at least.
Of course, if it wasn't off by 5%...
The bill was totally itemised, the total just wasn't the sum of all the numbers above it. I really wish I'd taken a picture of it now (I didn't really expect them to take it off me, I really expected them to explain that I was being a doofus) :(
In my experience this is usually because 'financial software' systems used to create invoices is sold by companies with 90% sales people, and maybe one or two developers. There seems to be little to no quality control. No one seems to care, since 'financial software' is very lucrative anyway.
In the beginning I tried to report the faulty invoices to the suppliers, thinking that the'd immediately press the big red emergency button and fix it, but in most cases the servicedesk employee does not care or even understand what I am talking about. Most of them send the 'thank you for your report, we are working very hard to fix the problem' email, but never actually fix it.
(In the UK, all prices shown in the menus are inclusive of VAT and then the total on the bill shows how much of the total is the VAT, so both the amount you can add up yourself from the menu and the line items shown on the bill should match the total you're paying perfectly).
Maybe I'm getting this all wrong, as I've never had to work with a register, nor do I have any experience with the software involved, but I hardly see the problem in accumulating a few cents of inaccuracy over time.
Anyway, yes, your main thesis right: it's entirely possible to do this with Javascript without screwing it up. But it wasn't in this one particular case.
I think the point is that the situation where floating point error changes the result when rounded to cents is rare. Which is true, except when it is not (e.g., if you happen to have a price that tickets commonly hit that results in an error.)
The problem is accumulating that inaccuracy when there are well-known, battle-tested libraries and data-types available specifically to avoid that problem. Except, maybe not as well-known as they ought to be...
People these days tend to have no idea how numbers are actually stored in bits and bytes. I'm more likely to find an English major who would be able to remember Big-Endian vs Little-Endian (albeit in the original Swiftian context) than a web developer who has any idea what that means. Forget about mantissas and exponents in floating point, or two's complement integers.
> if you do all your calculations in cents
That would certainly have been a good idea also, I bet.
With floating point, the assumption (x*y)/y = x does simply not hold.
I agree as much as the next person that you shouldn't represent money with floating point numbers, but I disagree that cash register software written in Javascript must automatically be incorrect (or more so that similar software in a different language). And I don't even like Javascript.
I'd assume that division is implemented as multiplying with the reciprocal (because thats faster). If that's correct, then any division by 3 (or any other number that is not of the form 2^i) breaks your cash register.
Because 1/3 represented as a floating point number equals 0.33333333.. and so on but not indefinitely, 3 * 1/3 = 0.99999999.. - which would be equal to 1 if you were using real numbers. But you don't.
"Buy one scroodad for a dollar, get two free."
"Three widgets for $2, limited time only!"
"$5 each, or three for $10"
They could have that £1.50 back if they could give me an explanation! =)
If you think about it, the item names are most likely coming from a database that just might not be in the right encoding (latin1 is still the default in MySQL I think). The symbols that do work are probably hard coded into the receipt's template, and hence don't have this problem.
Why a shop owner would store the price and currency symbol in an item's description is beyond me, but having worked in the POS world and seeing what shop owners do with their items I'd definitely believe it.
If you've worked in the POS world you know exactly why: bad software. It either doesn't support the use case the owner has or isn't easy to use.
Have you never watched Pulp Fiction?
https://www.youtube.com/watch?v=zoJAc_aSM7E
This video will explain exactly why a vendor might put a price in the description field.
Is there a reason for this? Or is it just another case of lol@mysql
Looks like I misremembered - 5 rather than 8 characters. But it isn't standard cp1252 and this can matter.
If it violates the documentation about invalid characters, that's a problem, but that's not latin1 being incompatible with 1252.
Plus that's a different scale of change because it's going from fixed width 7-bit to a variable width scheme.
That's probably a hint that it isn't the printers fault.
I would guess that some other system that is used to enter what's available on the menu is using CP 437 and somewhere an encoding step (CP 437 to Unicode) is missing so we get the ú character.
I wonder what character we would get if it was a "5€ cocktail" instead.
Bad collation.
In order to maintain backwards compatibility with existing
documents, the first 256 characters of Unicode are identical to
ISO 8859-1 (Latin 1).
This isn't true in a useful sense. It does look like it's true in Unicode codepoint space [1] but in any specific encoding of Unicode it can't be the case because latin1 uses all 0-255 byte values. For example, in utf8 it's only an exact overlap for bytes 0-127 (7 bit ascii).(Though maybe this means you could convert latin1 to utf-16 by interleaving null bytes with the latin1 bytes?)
[1] https://en.m.wikipedia.org/wiki/Latin-1_Supplement_(Unicode_...
Yes. In fact, things like JS JITs end up storing strings as either UTF-16 strings or Latin1 strings internally to take advantage of this fact.
And while the JavaScript APIs only allow you to deal with UCS-2, the string contents themselves are, in fact, usually UTF-16.
And of course it also works in the other direction. To safely read binary data into unicode strings just decode as latin-1 instead of utf8 and you won't run into validation errors since all byte sequences are valid in latin-1 while not all are in utf8.
Oh sweet summer child. No, ASCII itself was a problem. Before we had 8-bit character sets, we had 7-bit character sets:
https://en.wikipedia.org/wiki/ISO_646
This is why IRC considers [\] and {|} to be lowercase and corresponding uppercase letters respectively: it was made by a Scandinavian, and in their character sets, some accented characters occupy the same positions as ASCII [\]{|} would.
The story of character sets is the story of evolving common subsets: ISO 646 within ASCII, ASCII within “extended ASCII” (or at least, some variants thereof), Latin-1 within the Unicode BMP, the Unicode BMP within Unicode.
Oh and by the way, before we had 7-bit character sets, we had 6-bit (e.g. IBM BCD). And before those, we had 5-bit (e.g. Baudot code). And before that, we had different telegraph codes (variations of Morse code)…
The code that sends the "product name" was not, and doesn't correctly translate its input to the code page that the printer is using.
When I made a homemade POS system for a bar, years ago, I ran all the printers in bitmap mode and rendered the receipts in software, to sidestep this and other problems. The performance was still acceptable, but I think the reason many POS systems don't go this route is compatibility; they have to work with many models of printer and bitmap support is not universal, and even among those printers that support it I am not sure if it is standardised.
Cyrillic is not a language, it's an alphabet/script. Codepage 855 was used for Cyrillic mostly in IBM documentation. In Russia codepage 866 was adopted on DOS machines, because in codepage 855 characters were not ordered alphabetically.
>Even today, on modern windows machines, typing alt+163 will default to 437 and print ú.
It's only true for machines where so called "OEM codepage" is configured as codepage 437. But in Russia it's codepage 866 by default, so typing alt+163 prints г.
This is incorrect. It gives you an ANSI codepage. On old Windows version it would be a default ANSI codepage, on modern Windows it's a codepage associated with your input language. So if I type ALT+0163 with English keyboard layout I get £ from Win-1252, but the same combo after switching to Russian gives me Ј from Win-1251.
Entering numbers bigger than 255 just causes wraparound. For example, ALT+0835 also will give you £ instead of ₣.
>Everyone hates Microsoft.
If only it was only that... Microsoft has even worse encoding schemes. The ugliest I encountered was an "encoding" based on glyph indexes in ttf files.
Conversion is a pain in that case, and is uncertain... it also leads me to not so beautiful code...
https://raw.githubusercontent.com/kakwa/libemf2svg/master/in...
Even between Microsoft products (namely Office on Mac and Office on Windows), this scheme is not handled properly (the string is incorrectly handled as an UTF-16LE string on Office on Mac).
In Britain, it would be easy not to notice the incorrect symbol when setting up the machine. Elsewhere in Europe, it ought to get noticed quickly — but I occasionally get receipts in Denmark where the shop's address (or even name!) is corrupted, like "SkrÉdderi, LÎvstrÉde" instead of "Skrædderi, Løvstræde".
There are easy transliterations, but I input them on my British accounts to make a point. About half work correctly.
My first programming job was writing software for the MUMPS operating system on a DEC PDP-15, which had an 18-bit word size. PDP-15 MUMPS used 6-bit ASCII (which was uppercase only) because three characters fit nicely in an 18-bit word.
I'm willing to bet the problem here is that the descriptions of the items are stored in the database as ascii and not unicode.
I wrote the following article for Node.js to try and clarify the intersection of Unicode and filesystems, especially with regard to different normalization forms, and using normalization only for purposes of comparison:
https://nodejs.org/en/docs/guides/working-with-different-fil...
Even in DOS this caused issues:
1) NUL is same as space in cp437, but is all ones in many other DOS code pages. This causes strings output by some software written in C(++) to end in black rectangle (Notably in C++ version of Turbo Vision, including Turbo C++ IDE), background in many TUI applications consisting of thin 8px spaced lines is caused by same thing (see below for why it is rendered as thin lines)
2) DOS and language runtimes for DOS tend to ignore most control codes but still not all of them. In particular 0x07 BEL is useful character (often used as the dot in selected radio button), only way to get it on screen is to directly write into framebuffer.
3) MDA-style character generator (present on essentially anything but CGA) has special hardware logic for making cp437 box drawing characters one pixel wider. This means that all "right facing" box drawing characters have to reuse codes with this magic behavior and you cannot use these magic slots for normal characters that are wider than 7px. (And is reason for the thin 8px spaced lines)
So as part of this, and after years, I eventually realized the only way to make a scalable tool to lookup Han glyphs is to build upon UNIHAN: The Unicode Consortium's Han Unification effort.
I write about Unicode and UNIHAN in my own words here: http://unihan-etl.git-pull.com/en/latest/unihan.html
The challenge with Unicode and hanzi is there are many historical and regional variants to a single source Han grapheme of the same meaning.
So, each glyph or variant gets its own codepoint, or number, reserved. In fact, this years when Unicode 10.0 is cut, the new CJK Extension F will introduce 7,473 characters (http://unicode.org/versions/Unicode10.0.0/).
Thankfully, my only task is to make the database accessible in as friendly a way as possible. Which is actually a mammoth task, see, there are over 90 fields which are used to denote dictionary indices, regional IRG [1] indices (which are national-level workgroups that convene to add new characters), phonetics (mandarin, cantonese jyutping, and more).
The fields are dense. They pack in objects that are most easily split up by regular expressions. https://github.com/cihai/unihan-etl/blob/master/unihan_etl/e...
So a UNIHAN field for kHanyuPinyin (http://www.unicode.org/reports/tr38/#kHanyuPinyin):
U+5364 kHanyuPinyin 10093.130:xī,lǔ 74609.020:lǔ,xī
U+5EFE kHanyuPinyin 10513.110,10514.010,10514.020:gǒng
U+5364 is two values (separated by the space), then a list of items either of the colon (:), which are separated by commas.
You may wonder where this all comes from. The effort is global, but a good deal of it is thanks to people who took their time to contribute it, organizationally or personally. Take a look in the descriptions of the fields at http://www.unicode.org/reports/tr38/ for bibliographic info.
In any event, the hope is to create a successor to cjklib (https://pypi.python.org/pypi/cjklib) and have datasets for CJK available in datapackages (http://frictionlessdata.io/data-packages/). That way, sources of data are sustainable and not tied down to any one library.
[1] https://en.wikipedia.org/wiki/Ideographic_Rapporteur_Group
The printer probably use a default code page, and that's all. BTW, Unicode is not hard. The "hard" part is reading the device manual, and implementing encoding conversion properly. Also, in cases where no character selection is possible, in most cases you can use the printers in graphic mode.
* @15:32, "UTF-32 uses the same amount of bytes for (almost) all code points" — there is no "almost" about it; UTF-32 always uses 4 octets per code point.
* There was some amount of conflation between code points and characters.
* It was implied that len() will always give you length-in-code-points in Python 3, whereas it doesn't in Python 2. In Python < 3.3, it's code units (just like it is in Python 2), which on a narrow build will be 16-bit and thus wrong for strings w/ code points outside the BMP. This particular problem wasn't solved until 3.3 with the introduction of PEP-393.
The author's main points regarding the difference between text, and how you encode it, is good.
Unicode was started independently and later harmonized with UCS.
* The overwhelming majority of languages don't give you code-point level iteration over strings by default (and you probably want grapheme), most opting for code units — which is what an unsigned char ptr in C containing UTF-8 data will give you. (C++, Java, C#, Python < 3.3, and JavaScript all fall in this bucket)
* Linux, and most (all?) POSIX OSs store filenames as a sequence of bytes. What human chooses a sequence of bytes to "name" their files?
* Things like "how wide will this character display as in my terminal" are either impossible, or done with heuristics. Usually, it's not done at all; most DB CLIs I've used that output tabular data will corrupt the visual if any non-ASCII is output.
(Yes, some of this is in the name of "backwards compatibility".)
Saying for (let ch of str) in JavaScript iterates over the codepoints, not UCS-2 codepoints.
https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
When these absolute minimum intros talk just about encoding it misleads people to think that that's enough. I can't count the number of people who have read Joel's article and have the misconception that all user perceived characters are mapped to code points. I was one of those people. Just because ASCII and Latin-1 character sets can be mapped to code points does not mean that's how Unicode works.
At minimum every software developer must know four different levels:
* bytes,
* code points,
* combining character sequence
* grapheme clusters, extended grapheme clusters
Joel stops at the second level. He never gets into point where he explains how encode user perceived characters, how to detect grapheme cluster boundaries in the Unicode encoding.
examples: 각 , नी , நி
If a developer knows and understands the concept of character encoding ("It does not make sense to have a string without knowing what encoding it uses"), at least they will know how to read a string from one system and move it to a different system that expects a different encoding. They'll know that they need to call the relevant conversion routine in a specialized library that knows how to handle the conversion.
With this, maybe they won't be able to correctly build or modify their own strings directly. But being able to handle strings from an external system that produces them, and passing them to another system that consumes them, without breaking them in the process, IMHO does qualify as "the absolute minimum" they should know.
Joel gives the impression that he don't know that he don't know.
Knowing that you can't break a Unicode text string or insert text into the middle of Unicode string unless you know what language it uses is usually enough. They are just binary blocks you can't modify unless you have some extra info or uses specific libraries.
https://manishearth.github.io/blog/2017/01/15/breaking-our-l... gives an overview of most of the different things scripts do. Being aware of those helps a lot. It also gives a brief idea of how to deal with this stuff (usually it's just calling an API)
https://manishearth.github.io/blog/2017/01/14/stop-ascribing... gives an idea of how grapheme clusters work. You don't need to know the algorithm, just the stuff around it.
These things are things one can mostly ignore, until one can't.
Most input modes produce pre-composed (but not necessarily NFC) output. But some things will decompose (e.g., HFS+ will decompose filenames). So... if you cut-n-paste non-ASCII Unicode from an HFS+ file picker UI... you'll get into trouble if the software you paste into is unaware of these things.
Ultimately, every software developer needs:
- UTF validator (at least for UTF-8)
- UTF converters (unless only supporting UTF-8)
- case mapping (probably)
- normalization (almost certainly)
- collation (probably)
That's... not too bad.
Networking software )may_ also need:
- IDNA2008 implementation
- UTS#46 implementation
Word processing / typesetting software also absolutely needs to know about grapheme clusters in order to determine the size of each grapheme. Also: modern fonts.
https://eev.ee/blog/2015/09/12/dark-corners-of-unicode/
But one's sticking with only some part of Unicode support that they understand/need is easy, sure.
It's hard, because there's a lot more to learn and to do than if you stick to (say) ASCII and ignore the problems ASCII can't handle.
It's easy, because if you want to solve a sizable fraction of all the problems ASCII just gives up on, Unicode's remarkably simple.
In the eyes of a monoglot Brit who just wants the Latin alphabet and the pound sign, unicode probably seems like a lot of moving parts for such a simple goal.
Paraphrasing the joke about new standards: we had a problem, so we created a beatiful abstraction. Now we have more problems. One of the new problem being normalization.
It doesn't undermine the good that Unicode brought, but you can't say to have included some unilib.h and use its functions without understanding all the Unicode quirks and its encodings, because some of the parameters wouldn't even make sense to you, like the same normalization forms.
1. Either your restrict yourself to the kind of text CP437/MCS/ASCII can handle (to name the three codecs in the blog posting). In that case unicode normalisation is a noop, and you can use unicode without understanding all its quirks.
2. Or you don't restrict the input, in which case unicode may be hard, but using CP437/MCS/ASCII will be incomparably harder.
Not just it would be harder, but you couldn't get into space without it at all, so that got comparatively easier.
Is it still all easy, though?
I think Unicode is about as simple as it can possibly be given the complexity of human language, but that doesn't make it simple.
Also, a major reason why Unicode is large and complex is because languages and scripts are large and complex. Unless we all agree on using simple computer-friendly languages and scripts that complexity is not going to change, and the need of working with older scripts (e.g. for historians and researchers) still requires something like Unicode. Unicode is the kind of thing that emerges from a messy world, and unsurprisingly it's messy as well.
For example, in olden times, or when restricted to ASCII, the Nordic letter "å" is written "aa", but it is still sorted at the end of the alphabet — "Aarhus" will be close to the end of a list of towns.
In Welsh there are several digraphs, single letters written with two symbols. The town "Llanelli" has 6 letters in Welsh. (There are ligatures, but I don't think they're often used: Ỻaneỻi.)
Now it looks like we're up to our old tower building ways again, except this time with computers and data. So God smirked and gave us Unicode.