Character Encodings For Modern Programmers
blog.gatunka.com
blog.gatunka.com
We no longer mess around with code pages, or old-school multibyte encodings like Shift-JIS, and we also don't compile Windows executables in "UNICODE mode", instead we keep all strings as UTF-8, and convert from and to UTF-16 (on Windows) or UTF-32 (on UNIX-like operating systems). Conversion to and from UTF-32 is hardly necessary though, since outside Windows, everything seems to use UTF-8 anyway.
Instead of using OS functions like MultiByteToWideChar() or the iconv library, we have integrated the LLVM/Unicode standalone UTF conversion functions: http://llvm.org/docs/doxygen/html/ConvertUTF_8h_source.html, although with C++11 these conversions are now builtin.
A few other notable points: - properly handling IME input for Asian languages can be tricky (for fullscreen 3D games)
- Arabic text rendering (not because of right-to-left, but because the character appearance changes depending on whether a character is at the start or end of a word, and there is nearly no sample code around which demonstrates this behaviour (and the one we found had all Arabic comments)
- some Asian languages require incredibly huge font textures (most 3D-game text renderers are font-texture based as far as I'm aware of)
[edit: removed redundant link]
- a few code pages (such as CP864 DOS Arabic) use the arabic percent sign ٪ (0x066A) at codepoint 0x25 instead of the standard ANSI %.
- Mozilla Firefox and Thunderbird approximate the Arabic codepages (including ASMO-708) to ISO-8869-6, which breaks XLS parsing from files generated from certain international versions of Excel (which actually forced me to build a library for the various character encodings: https://github.com/SheetJS/js-codepage)
- We went from 128 ASCII characters to 110,187 currently-assigned Unicode characters (or Encoded Character code points, to use precise language)[1]. That's a thousand-fold increase. Who would have thought we'd need so many!
- Unicode can represent 1,114,112 code points. So there's lots of room for yet more.
[1] http://babelstone.blogspot.ca/2005/11/how-many-unicode-chara...
Keep in mind that Unicode only emerged in the 1990s and Excel predates unicode by about a decade.
http://www.joelonsoftware.com/articles/Unicode.html (The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets)
and then return to read the featured article.
I'll observe that this article is very C-centric. If Python is how you think about things, I recommend the Python Unicode HOWTO, which of course has two different relevant versions at the moment:
For Python 2: https://docs.python.org/2.7/howto/unicode.html
For Python 3: https://docs.python.org/3/howto/unicode.html
If you're bored or not feeling modern today, read http://en.wikipedia.org/wiki/EBCDIC
For practical Python advice (I'm sure this was on hacker news) http://lucumr.pocoo.org/2013/7/2/the-updated-guide-to-unicod... is a great read.
You can still encounter these in printed receipts from old POS systems. The half-width katakana still exist as separate codepoints in Unicode, and are quite popular in emoji and such.