Projecting Unicode to ASCII
johndcook.com
johndcook.com
>>> import unicodedata
>>> print( unicodedata.normalize('NFKD', "éèêàùçÇ").encode('ascii','ignore'))
eeeaucC
But if you need to project (transliterate to ascii) Arabic, Russian or Chinese, unidecode is close to black magic: >>> from unidecode import unidecode
>>> unidecode("北亰")
'Bei Jing '
Anyway, always remember that str.encode(), str.decode(), open() and many other related callables have an "errors" parameters that allow you to deal with unkown solutions when encoding or decoding: >>> print("Père Noël".encode("ascii", errors="ignore"))
b'Pre Nol'
>>> print("Père Noël".encode("ascii", errors="replace"))
b'P?re No?l'
I'll conclude with the mandatory "use Python 3" (3.7 if you can, it has many utf8 fixes: https://vstinner.github.io/posix-locale.html), since you'll be in a world of pain if you deal with non-ascii in Python 2, and EOL is next year :) Tic, Toc...DIN 5007 Var. 2 specifies that for the purposes of sorting, ö is replaced with oe. This also applies to ä (ae), ü (ue) and ß (ss). This same replacement rule is also used on passports and IDs.
Unidecode does not handle this correctly.
The unidecode author wrote about this:
> In German, there's the typographical convention that an umlaut (the double-dots on: ä ö ü) can be written as an "-e", like with "Schön" becoming "Schoen". But Unidecode doesn't do that-- I have Unidecode simply drop the umlaut accent and give back "Schon".
> (I chose this not because I'm a big meanie, but because generally changing "ü" to "ue" is disastrous for all text that's not in German. Finnish "Hyvää päivää" would turn into "Hyvaeae paeivaeae". And I discourage you from being yet another German who emails me, trying to impel me to consider a typographical nicety of German to be more important than all other languages.)
Every such projection is a lossy projection and is expected to possibly generate in an ambiguous result that must be interpreted through context. Hence the strange but comprehensible ,,Godel'' and ,,Malmo''. Cook is not claiming that his code replaces the need for Unicode!
(There is a minor linguistic irony that the origin of the umlaut was scribe's shorthand when an E vowel inflection was turned into tiny E written above a letter, almost like a ligature, which became a pair of dots. But that character was then used in other languages differently, much as a loanword from a different language usually changes its meaning in the new language).
That might be a dangerous assumption. The project page explicitly states:
> Transliteration of languages like Chinese is a very complex issue and this library does not even attempt to address it. It draws the line at context-free character-by-character mapping.
I.e. it is not black magic; just a mapping of Unicode characters to static ASCII transliterations. The results will certainly be incorrect in some contexts.
Unidecode doesn't do language-specific transliteration and really works best for user-invisible things, like database identifiers or normalization.
CJK characters in particular are very problematic, since they must be transliterated differently depending on the locale. Over the years I have received many angry mails from people that were deeply offended by an error in transliteration they saw in an URL or something.
Unihandecode is a fork of Unidecode that tries to address this:
Time has devoured my comment on that post, which extended the transliteration table to Eastern European languages and proposed to use mnemonic names like int(u'\N{Latin capital letter AE}') instead of 0xc6.
https://github.com/pudo/normality/blob/master/normality/tran...
Pingtype tries to solve all these problems. If there's a need for it to be ported to Python/etc then I'd be happy to do so!
Not sure how well it handles CJKV chars though.
$ echo 北亰 | iconv -t ASCII//TRANSLIT
?? uconv -x ':: Any-Latin; :: Latin-ASCII; [:^ASCII:] > \_'
Explanation: Any-Latin
Transliterates from non-latin scripts to latin script (e.g. "γραφὴν" --> "graphḕn"). Latin-ASCII
Tries to asciify characters as much as possible by discarding accents, splitting ligatures, replacing Unicode quotes with regular quotes, © --> (C), etc. [:^ASCII:] > \_
Replaces any remaining non-ASCII characters with an underscore.My understanding (which could be quite wrong) is that Windows refers NFC and MacOS prefers NFD (as do I but I understand the NFC desire) but that the filesystems themselves do not do normalization. In which case, modulo delimiters, every filename in one is legit in the other.
My suggestion of utf8 was to make a byte-order and code point width invariant form that should be completely reversible.
But if that doesn't work, never mind.
>Each transform rule consists of two colons followed by a transform name.
http://userguide.icu-project.org/transforms/general#TOC-Comp...