The Absolute Minimum Every Software Developer Must Know About Unicode
joelonsoftware.com
joelonsoftware.com
* Use UTF-8 for external text, whenever possible. If your collaborators have other ideas, bribe them with tasty cookies or something, because this right here solves a lot of hassles. There are some circumstances in which a different encoding might have its advantages, but it is tremendously reassuring to be able to say "Ah, text! I shall decode it as UTF-8!" and be right. This has the advantage of being compatible with ASCII input, and avoiding the perennial UTF16/UCS2 confusion.
* Make it explicit that you're using UTF-8. For example, if you're making a web page, be sure to set "Content-Type: text/html; charset=utf-8" in the HTTP headers, to make the browser's content encoding detection trivially correct.
* When dealing with strings in your favorite programming language, always know whether it's an array of Unicode code-points, or of bytes in UTF-8, or some third messed-up thing. Not all strings are the same kind of thing! Unless your programming language has this distinction enforced by its type system, of course.
* Be aware when you're crossing the boundary between Unicode code points and Unicode in some external encoding. Decoding can fail, so be prepared. It's best to reject invalid text as early as possible.
* When in doubt, use other people's code for Unicode handling. For most of the crazy crap you run into in the wild, there is well-tested crazy-crap-handling code.
And really. If a developer doesn't understand everything regarding Unicode, tell him it's simply mandatory to use UTF-8. Other encodings are for people who really really know what they're doing.
What is an 'array of Unicode code-points'. I actually reread the article thinking that my unicode chops were getting rusty if I didn't know what that meant, but after rereading it, I still don't know what you mean.
[I wrote a half a dozen 'do you mean _____' suggestions, but I think I'll just let you explain. :) ]
This is how way too many people think that UTF-16 works: each code point gets 16 bits, and you have an array of them, so you can count characters, do O(1) random indexing, and so on. This is a harmful myth, of course. Code points do not correspond neatly to glyphs, and UTF-16 is a variable-width encoding, although most people don't use code points outside the basic multilingual plane, so a lot of people can get away with pretending that it's fixed-width, until they can't.
The most maddening instance of this confusion that I've seen so far is in Python's Unicode string handling. Guess what happens when you run this Python code to find the length of a string containing a single Unicode code-point:
print len(u"\U0001d11e")
This will print either 1 or 2, depending on what flags the Python interpreter was compiled with! If it was compiled one way (the default on Mac OS X), then it uses an internal string representation that it sometimes treats as UTF-16 and sometimes as UCS-2. With another set of flags (default on Ubuntu, IIRC) it will use UCS-4 and do the Right Thing. For the same task, Java gets the string length right, but requires you to explicitly write the string as a UTF-16 surrogate pair: "\uD834\uDD1E".The redeeming virtue that both share is that they will do the right thing if you treat everything as variable-width encoded, use the provided methods for encoding and decoding, and avoid the hairy parts left over from when people naively assumed that UTF-16 and UCS-2 were the same thing and that they ought to be enough for anybody.
I tried to convince Fog Creek to abandon the obsolete, non-standard and proprietary 8-bit Windows-CP1252 character encoding in the E-mails that FogBugz sends. They refused, reason given: "joelonsoftware.com is a blog. Fog Creek Software is a business."
I think this could be because you use messagelabs.com to send out the emails.
Try including the funny character of your choice in an email. It should show up just fine.
I think this is a very good compromise.
Received: from mxny1.fogcreek.com (mxp11.fogcreek.com [64.34.80.172]) by […]
Received: from hbny4 (hbny4.fogcreek.local [10.2.0.48]) by mxny1.fogcreek.com (Postfix) with ESMTP id 2386295B33 for […]; Tue, 10 Jan 2012 19:18:45 -0500 (EST)
Priority: normal
Mime-Version: 1.0
Content-Type: text/plain; charset=windows-1250
Content-Transfer-Encoding: quoted-printable
Message-Id: <20120111001845.2386295B33@mxny1.fogcreek.com>(Some old Mac/Linux browsers resisted doing so and would show 'moronic' private 1252 characters as missing characters, even if a correct font was installed.)
Fast forward 8 years... I've joined Stack Overflow about a week ago and my first accepted answer was about Unicode and string handling in Python 2.x. Just yesterday I was thinking I should reread this exact post, to refresh a few points and keep shooting fish in that barrel. I guess I should be grateful Python 3 has gone full-Unicode... except that now people ask how to emit ASCII with it. And the Python community is among the most clued-up on the subject (probably as a reaction to how bad it was handled in 2.x).
And I'm not even a fucking software developer.
I wish this was the case. And for 99% of you it might be. But there's still lots of EBCDIC (and antique COBOL) out in the wild that increasingly has to interact with the real world. One of the chunks of telephony software I'm responsible for has to interface with an EBCDIC system, requiring some basic translation to UTF-16. (And then eventually has to pass some of this now UTF data back to the EBCDIC system.) Not exactly difficult until you get to the various numeric encodings and no one can decide if they want Binary-Coded Decimal, Pic9s, pure int, signs, which endianess (if they even know about endianess), et cetera.
EBCDIC wasn't dead in 2003 when Joel wrote this and it certainly (and unfortunately) isn't dead in 2012.
Fortunately, the recode program makes it easy to switch between different encodings.
I wouldn't even think about reading it unless I saw it posted on HN. However as long as I've got to deal with character encoding problems like this:
http://ibm-china.jobs/branch-admin-lan-zhou/jobs-in/
I'll still reap some benefit out of re-reading this essay.
Written in 2003, 7 years after UCS-2 was obsoleted by UTF-16 because Unicode 2.0 was too big for just 16 bits. UCS-2 is not UTF-16. Windows NT wasn't UTF-16 either, IIRC.
Of course, Microsoft kept telling everyone that 16bits-per-character is Unicode.
Just use UTF-8.
If you do a lot of string manipulations, you're better off with either UTF-32 or dumbified (16-bit fixed) UTF-16, otherwise you will have to count characters from the beginning of the string every time you need to access the nth character within the string. Moreover, if you deal with a text with a lot of characters between 0x0800 and 0xFFFF (e.g. East-Asian languages) you're much better off with UTF-16 as you will save a whole byte per character.
As for East Asian text, you have a point: it will usually be shorter in UTF-16 than UTF-8. Before making this decision, though, ask yourself how much that extra space is worth to you. Is it worth dealing with possible encoding hassles? (The answer to this may be yes, but it's a question that should be asked.) Also, on a lot of data, there are many characters from the ASCII range mixed in with the East Asian text. I did an experiment a while back where I downloaded some random web pages in Chinese, Japanese, Korean, and Farsi, and compared their size in UTF-8 and UTF-16. Because of the amount of those documents that was HTML tags, all four pages ended up smaller in UTF-8.
Another example: for a lex-and-yacc type of parser, you can use regular expressions to split a string into tokens, and then use a parser on that stream-of-tokens representation. None of this requires character indexing; just byte indexing.
>> In UTF-8, every code point from 0-127 is stored in a single byte. Only code points 128 and above are stored using 2, 3, in fact, up to 6 bytes.
In Wikipedia:
>> UTF-8 encodes each of the 1,112,064[7] code points in the Unicode character set using one to four 8-bit bytes
Reference: http://en.wikipedia.org/wiki/UTF-8
e.g. http://golang.org/pkg/utf8/
>> UTFMax = 4 // maximum number of bytes of a UTF-8 encoded Unicode character.
Also 4 is max. value used in MySQL server.
Say a UTF8 string is ae 31 c1 12.
Now how do we decide whether it has the characters "31","c1","ae","12" or the characters are "ae 31" and "c1 12" or even "ae","31 c1" and "12".??
EDIT: Never mind!..found my answer here http://stackoverflow.com/questions/1543613/how-does-utf-8-va...
http://news.ycombinator.com/item?id=1219065
(I do not work for cafepress)
Also, if you write software for money, ignoring Unicode is just plain incompetent. I can't count the number of times I've had packages shipped to "Biały Kamień" or "Bia&#322;y Kamie&#324;" street instead of "Biały Kamień". If you expose even a single name or address field in your software, you need to handle Unicode.
And yes, you need to handle Unicode even if you want to limit yourself only to the US market. Your customer might have an umlaut or an accent in his name.
I agree and I should know better. I even opened the link. But after reading for 10 seconds, I closed it. I just wasn't motivated enough. Even had to fix bugs and issues related to this just last month. But every time I do, I just go and find out enough to solve the problem and then never really dig deeper. Don't really know why it is this way.
I cared about it. I will forget to eat and sleep if I am studying something I am motivated about. The rest I just learn as needed only if am forced to (read "when stuff breaks"). Not a very good approach I guess.
Don't be so sure that it isn't broken.
Actually you are right. Most input and output of our software is not text. And we mostly work for the US military (so UI is in English). This probably explains why. However in general I feel I should know more about. It is sort one of those things like when we were talking about interview questions and someone brought up Pascal's triangle. I felt I should have known about what it is, but I couldn't remember. It is not something I need to know for work, but rather something I felt embarrassed for not knowing.
For instance, when I search on Facebook for names, I never type accented e-s, but it stills brings up my friends who have accented e-s in their name.
http://www.tbray.org/ongoing/When/200x/2003/04/06/Unicode
Tim has spent more time working with bodies of text [1][2] than many of us, so his perspective is very useful and the article is thoughtful.
[1] Oxford English Dictionary [2] XML committee
I find out that even working as a user with text that is right-to-left in places and in other places left-to-right is hard. It's quite easy to deal with it programatically, but only before you have to display it somehow.
Ok, sorry my ramblings.
http://en.wikipedia.org/wiki/Mathematical_operators_and_symb...
square square square
Contrast with Firefox
"When I discovered that the popular web development tool PHP has almost complete ignorance of character encoding issues, blithely using 8 bits for characters, making it darn near impossible to develop good international web applications, I thought, enough is enough."
So that hasn't changed.
I'm not saying PHP's UTF-8 handling is great by any means, but the claim was that it's "nearly impossible." I'm suggesting that one should instead say "Building a UTF-8 compliant site in PHP is annoying, and requires more work than one would prefer, but if you do a bit of research, it's not that hard."
Because if you switch encoding you need to modify the code. All languages actually supporting utf-8 use the very same functions whatever the encoding is, eventually you simply need to declare that you're using utf-8 but that's all.