Known anomalies in Unicode character names (2017)
unicode.org
unicode.org
Personal example.
In Georgian (a few millions native speakers) we had historically three different independent alphabets: mtavruli, nuskhuri and mkhedruli. All represent same letters, but different ways to write them. Alphabets are represented in Unicode under different code points. Only mkhedruli is used currently (last few centuries) and therefore most fonts have glyphs for mkhedruli only.
Also there is no concept of capitalization in Georgian, there is no "a" and "A", we have only "ა", one does not start a sentence, a word or a personal name with some variant of a letter. Anna is ანნა in Georgian, the first and the last letters are written the same way.
Now, because of some reasons, letters of mtavruli alphabet are registered as uppercase versions of letters of mkhedruli alphabet. Mind blown.
And since well written software uses rules defined by Unicode standard, via ICU library or some platform specific routines, many desktop and mobile applications as well as major web-sites are hard to use. First letters of sentences are capitalized, but since mtavruli glyphs are missing from fonts, they are rendered as ⍰.
So instead of
მაგათი დედაც ვატირე.
we see
⍰აგათი დედაც ვატირე.
and have to guess what was the first letter. But who cares, if it is not used in U.S. right?
Definitely. For Georgian alphabets that were one of the original scripts included in Unicode 1.0, as an example, the reference fonts were designed by Georgian type designers Anton Dumbadze and Irakli Garibashvili. If something was very off they must have noticed that.
And it seems that the Georgians themselves were the ones who pushed for this change: https://unicode.org/L2/L2016/16034-n4707-georgian.pdf. It seems this isn't a case of "ignorant Westerners screwing up Unicode" but "ignorant programmers who don't understand how Unicode works screwing it up," as noted in the sibling comment.
We have English speaking programmers swearing by ASCII despite it not having all the characters needed for English (I can't find it now but there was a report by the Library of Congress that came to that conclusion). I suspect their Georgian counterparts have similar attitudes towards text.
I've heard this argument a few times, and each time it's just as contrived, particularly if you stick to American English. While I'm happy to acknowledge it's not good enough for other languages, ASCII is plenty good for English. Unless you're trying to type cafe with an accented e to be fancy or type that old conjoined-ae character, ASCII is fine for everyday usage. See the fact that most Americans have never installed a separate keyboard layout.
It’s good enough for the style of American English adopted as a concession to typewriter limitations, since that’s pretty much exactly what ASCII is designed to support. OTOH, American English as used outside of that context doesn’t use straight quotes, but rounded ones in each direction with different meaning, uses various diacritical marks, and also uses the cent sign (¢), the obelisk (†), diesis (‡), pilcrow (¶), section sign (§), division sign (÷), multiplication sign (×), less-(and greater-)than-or-equals sign (≤ and ≥), angle brackets (⟨ and ⟩, which aren’t the same as less-than/greater-than symbols), the interrobang (‽), and—the next three distinct from the hyphen—the minus sign, en-dash, and em-dash.
> See the fact that most Americans have never installed a separate keyboard layout.
Common software used for text, like MS-Word, provides convenience methods for many non-ASCII characters without an keyboard layout change.
It means that because the English language uses punctuation and other non-alphabetic symbols, not just letters of the alphabet, as an examination of any substantial corpus of written material in the language not solely consisting of material where the writer was constrained by ASCII (and, heck, even those where they were, though the set of non-alphabetic symbols used would be narrower) would instantly reveal.
> Half you quoted are general math symbols that have nothing to do with English.
5 (7 if you mistakenly include angle brackets, which do have textual, non-mathematical uses) out of 15 is not half, but mathematical symbols are used in written English.
If my weird squiggley that I use in my specific niche field isn't represented by any Unicode code point, does that mean that Unicode isn't good for representing English?
Regardless of the technical arguments, from the perspective of non-programmers, when they see things like the CA DMV being incapabale of doing things as simple as recording a diacritic over a person's name, it doesn't reflect well on the software industry and programmers as a professsion.
While true, I think the overwhelming majority of German-Americans anglicized their names several generations ago. For instance, my great-great grandfather switched from Müller to Miller. This is probably less true of more recent arrivals.
https://www.merriam-webster.com/dictionary/cafe‘: “variants: or less commonly cafe”
https://dictionary.cambridge.org/dictionary/essential-americ...: “noun (also cafe)”
https://www.dictionary.com/browse/cafe also prefers to see the accent
That spelling preference may be shifting. I couldn’t find evidence for it, though.
Really? The data say otherwise: https://books.google.com/ngrams/graph?content=caf%C3%A9%2Cca...
Do you have another source that supports the popularity of the diacritic variant? Rather, café doesn't increase in popularity until after 2000, and thus would appear to be an affectation of sophistication.
Windows-1252 encoding was supported on American computers long before Unicode. I remember “smart quotes” being promoted as a feature of word processors.
Even for people who didn't care for the quotes:
• everyone appreciates ¢, ° for °F or 46° N, ¼, ½, ¾, × and ÷,
• people handling large texts want §, †, ‡ and maybe ¶,
• businesses want ©, ® and ™,
• scientists and engineers need µ, ±, ², ³, ·.
Web browsers have made it more difficult for the average user to enter these characters, but software like Microsoft Word has supported easy ways to insert many of them automatically.
Example: https://upload.wikimedia.org/wikipedia/commons/7/79/Bill_of_...
The most common English spelling is Kiev. That reflects the English pronunciation. (The British one anyway.)
But many significant European cities with a long history have a different name in English than the local languages: Cologne, Munich, Florence, Naples, Warsaw, Copenhagen, Gothenburg, Prague.
In French: Douvres, Londres, Édimbourg.
In German: Edinburg.
In Italian: Edimburgo, Glascovia, Londra, Novocastro, Dublino.
In Chinese: Jiànqiáo (剑桥) (lit. "sword bridge", Cambridge)
Some people just don't care.
ToUpperCase("dz") => "DZ"
ToTitleCase("dz") => "Dz"I routinely see code that treats titlecase as “set individual character in this Unicode string to its upper case equivalent”... It wouldn’t surprise me at all to find out this kind of titlecase vs uppercase mistake is a very widespread issue.
If you press "s" key it types "ს" (S), but if you press "shift+s" it types "შ" (SH), which is entirely different letter. Georgian alphabet has 33 letters, 7 more letters than Latin, so letters ჭღშჟძჩთ are mapped to shift+something keystrokes.
Russian language has 33 letters too, but Russian layouts override 7 punctuation keys ,.;'[]~ so pressing "," will not type a comma, but a letter "б", pressing "]" will type "ъ", so Russian layouts remap punctuation symbols to shift+number keystrokes. Georgian layouts keep punctuation key assignments same as for English and use shift+letter for additional letters.
Modern fonts may have mtavruli styled glyphs, that is true, but they are placed on mkhedruli code points anyway, which is exactly the problem, no widely used fonts support this case switch anyway and render ⍰.
Mtavruli has already been supported by the following fonts: Google Noto (Android 11 DP, Arch Linux), SegoeUI, Calibri (Windows 2019 update), Helvetica Neue (macOS Catalina, iOS 13).
This isn't intended to sound rude, but why should an American care? Dismissing the argument that for better or worse many people hold doesn't do any good. Making it might convince them otherwise.
After all, the POSIX standard includes slash as a name for the character, and that an ISO standard: https://pubs.opengroup.org/onlinepubs/9699919799/ What's more, Unicode refers to "\" as "backslash".
Very weird.
Apparently Unicode has renamed it be "REVERSE SOLIDUS" while still tolerating "BACKSLASH" as a name for it.
I don't expect anyone to ever use the words "REVERSE SOLIDUS" in my presence.
Interesting, it entered computer use in the era of Algol so they could write their /\ and \/ operators.
I'd be happy if nobody ever referred to the word "backslash."
I heard it read in a URL in a TV commercial this week. I can't believe it. People have been getting that wrong since AOL days.
Yes, there are people who get things wrong on the internet. Just gently correct them and move on. I don't think trying to change to other terms like solidus and reverse solidis are going to make it any better.
And I like :// being in URLs. It makes them incredibly easy to disambiguate them from anything else.
See also, ghost kanji: https://www.japantimes.co.jp/life/2018/10/29/language/ghost-...
The vast majority of CJK characters have systematic names like "CJK UNIFIED IDEOGRAPH-72AC" which aren't really subject to errors in the same way as the more verbose names used for other scripts.
There are some pretty clear patterns to the character ordering, though -- if you look closely, you can see big runs of characters which share radicals. For example, characters 5000 through 500F are:
倀 倁 倂 倃 倄 倅 倆 倇 倈 倉 倊 個 倌 倍 倎 倏
all of which have the 亻radical on the left side. It's clearly not arbitrary.
That's because those runs were allocated at once (e.g. CJK Unified Ideographs Extension G spanning from U+30000 to U+3134A) and they were systematically ordered by radicals and stroke counts. Note that radicals and stroke counts are fairly arbitrary and can differ among character sources (the Unihan database has a ton of them). While fairly predictable, this ordering is ultimately arbitrary.
Except for 倉 it seems.
For many scripts we can come up with some set of names that are suitable for identification. Han characters, among others, aren't. One can say that Han "characters" are identified by its shape, but a line between two differently perceived characters is extremely unclear. Some may recognize two characters as same, some not, some would even try to add a stroke or so to differentiate the character. Han characters are thus identified by providing multiple properties for them (the Unihan database), and the character names are just placeholders.
The one with Lao letters "LO LING" and "LO LOOT" is actually very blunt error -- names for them are swapped. It's pretty safe to assume that whoever was making the standard does not know this writing system.
> U+262B FARSI SYMBOL
> This symbol is so named because as symbol of Iran it cannot be encoded in ISO standards.
It would seem strange that ISO standards cannot refer to countries, so I'm guessing this is something specific to Iran?
<quote>
As noted by Roozbeh Pournader:
Neither Farsi, nor a symbol. In real life, it is the official emblem of the goverment of the Islamic Republic of Iran.
Technically that would make it a logo and thus not a suitable candidate for encoding. But Roozbeh also noted:
Exactly. The funny fact is that it has been in Unicode since 1.0...
</quote>
How reliable this is I don't know, but it sounds plausible.
[1] http://archives.miloush.net/michkap/archive/2005/01/29/36320...
Aside: Microsoft sucks for having purged blogs like that one. The archive you're looking at captures most of what was once Michael Kaplan's blog at Microsoft. Kaplan was terminated by Microsoft and then died, and his blog was one of a large number of valuable blogs with insights into Windows technologies that at some point were "tidied away" because they didn't fit whatever nonsense brand vision somebody had that week.
https://aka.ms/RootCert will always be Microsoft's Root Trust programme documentation (how a Certificate Authority like Let's Encrypt gets themselves listed as trusted in Windows) even when Microsoft decides that page should now be in the form of a GIF anim or a 3D bullet hell game.
I suspect this is actually shorthand for "ISO wanted to avoid referring to countries in Unicode character names". But that ship has clearly sailed nowadays with characters like U+1F5FE ("SILHOUETTE OF JAPAN").
This may sound pedantic, but Japan the island or Japan the nation? It seems to refer to the island(s). On the other hand, the "farsi symbol" is of a particular nation rather than a given set of boundaries. Nations can be renamed, but the island of Japan refers to the chunk of earth. Add to that that nations get redrawn every so often in the middle east and I can understand the position.
I wouldn’t normally pick something up like this, but isn’t the whole point of ASCII to have a way to represent non-English characters?
That's funny. Hacek is a Czech invention (it was actually a dot originally), but I always assumed that "caron" is the official "correct" (english) name.
—
Also, from the parent page for that link:
“ Unicode Technical Notes provide information on a variety of topics related to Unicode and internationalization technologies.
These technical notes are independent publications, not approved by any of the Unicode Technical Committees, nor are they part of the Unicode Standard or any other Unicode specification. Publication does not imply endorsement by the Unicode Consortium in any way. These documents are not subject to the Unicode Patent Policy.”
SOURCE: https://unicode.org/notes/