Unicode In Python, Completely Demystified
farmdev.com
farmdev.com
One other unicode library I can recommend is icu, Python bindings to IBM's ICU. It solved some Turkish specific problems I have. (Turkish alphabet has I,İ, ı, i which makes upper-lower case conversions tricky. For example, mayıs becomes MAYIS when capitalized while ENGLISH becomes englısh when lowercased.)
Read http://en.wikipedia.org/wiki/Turkish_dotted_and_dotless_I , http://www.joelonsoftware.com/articles/Unicode.html and http://www.codinghorror.com/blog/2008/03/whats-wrong-with-tu... for more info.
Another issue, if you are doing web dev is, how to represent non-ascii characters in urls. One choice is url encoding, the other is slugifying the url, that is, choosing an ascii equivalent for the non-ascii character. For example, Django has a slugify function that helps with this. It converts, for example, über to uber.
What does it convert 'ぬびばざべ' into?
キャンパス -> kyanpasu
Αλφαβητικός Κατάλογος -> Alphabētikós Katálogos
биологическом -> biologichyeskom
From: http://userguide.icu-project.org/transforms/generalEdit: reread, OP was talking about Django transliteration, which is much simpler.
Spolsky's article on character encoding is equally as good. http://www.joelonsoftware.com/articles/Unicode.html
I recommend both to every programmer.
BTW using `mbcs` encoding on Windows is more compatible than ASCII in most cases.
Then there is no reason Python using ASCII to handle 8bit byte and yield error.
When I last looked at this, Python doesn't attempt to deal with the Windows console code page.
For example, freshly installed python3.2 from macports:
>>> len(chr(119074))
2
>>> chr(119074)
'𝄢'
>>> print chr(119074)
𝄢
(Surrogates are so much fun).Python 2.7.1 compiled with UCS4:
>>> len(unichr(119074))
1
>>> unichr(119074)
u'\U0001d122'
>>> print(chr(119074))
𝄢
Notice the capital U takes 8 instead of 4 hexidecimals.Someone should do a comparative analysis across languages. My guess is that (in-browser) javascript is probably one of the few languages to not have significant unicode problems, although there are surely some.
I mean, how dumb does the system have to be to reject the @ char in the password field? Or worse yet, accept only _numbers_, as my university does.
I think we did something terribly wrong somewhere in the char set conventions.
http://benlynn.blogspot.com/2011/02/utf-8-good-utf-16-bad.ht...
http://stackoverflow.com/questions/1049947/should-utf-16-be-...