UTF-8 good, UTF-16 bad
benlynn.blogspot.com
benlynn.blogspot.com
"UTF-16 has no practical advantages over UTF-8" - UTF-16 is better - i.e., quite a bit more compact, since, e.g., Chinese characters always take 3 characters in UTF-8, for asian languages.
"A related drawback is the loss of self-synchronization at the byte level." - Maybe this is a problem, maybe not. Maybe the failure of UTF-8 to be self-synchronising at the 4-bit level is a problem is some circumstances. I don't mean to be flippant, but the wider point is that with UTF-16, you really need to commit to 16-bit char width.
"The encoding is inflexible" - I think the author has confused the fixed-width UCS-2 and the variable-width UTF-16.
"We lose backwards compatibility with code treating NUL as a string terminator." - Not true. NUL is a 16-bit character in UTF-16. Use a C compiler that supports 16-bit char width.
Yea, it really should be "We lose backwards compatibility with code treating a zero 8-bit byte as a string terminator."
>"UTF-16 has no practical advantages over UTF-8" - UTF-16 is better - i.e., quite a bit more compact, since, e.g., Chinese characters always take 3 characters in UTF-8, for asian languages.
For pure CJK text files, yes.
"I think the author has confused the fixed-length UCS-2 and the variable length UTF-16." - No, in the same sentence you are quoting he mentions that utf-16 is limited to 0x110000 codepoints, contra utf-8 that is specified to expand up to 6 bytes.
"Not true. NUL is a 16-bit character in UTF-16. Use a C compiler that supports 16-bit char width." - I won't be converting my source code into utf-16 any time soon. Besides, C is not a good example where this is an actual problem, the "unicode strings" will still just be represented as binary blobs - different depending on which encoding you choose. It's more of a problem in e.g. Python, where the native string is an array of characters, not just a null terminated blob. I wouldn't encode my source code in utf-8, however, if my editor my default happened to support utf-8, this wouldn't be an issue unless I enter an non-ascii compatible code point. I see this as a big plus.
"supports 16-bit char width" - I mean the compiler has CHAR_BIT == 16; this is independent of the character encoding of the source.
In some contexts that may matter. In other cases, you can expect enough ASCII mixed in to outweigh that effect. For example, on the Web though, if you're sending HTML, the savings from using 8 bits for the ASCII tags will nearly outweigh the cost of extra bytes for content text. Gzip shrinks that further to essentially no difference.
EDIT: oh, and before any "semantic-nerd" comes along: I am fully aware that 0xXXXX are two bytes, so, if you want, read "two-byte" for every time I mention "byte" above... (doh ;))
I don't see how this gives any advantage to either of the encodings. Both for UTF-8 and UTF-16 you have to implement some decoding to reliably count the number of characters.
Thinking about it, I don't know what codepage the UTF-8 128-255 code points map too, if any, though; could you explain? If you treat UTF-8 as ASCII data (as one byte, one character, basically), does it generally work with chars in the [127-255] range.
If you can guarantee there are no bytes with the high bit set (a simple byte mask comparison which can be done 32 bits at a time for speed), then UTF-8 devolves to ASCII and you can calculate character-length without any division at all!
Counting the number of characters in a string is a misleading benchmark, because it's not a thing that people need to do terribly often - if you're rendering text to a display, "character count" is not enough, you need to know things like font metrics and which characters are combining, zero-width, or double-width. Concatenating strings is easy, and only cares about byte length, not character length. And if you really, really find character-length calculations on your hot-path, you can always just store length-prefixed strings instead of terminated strings.
And the "ASCII argument" is quite antiquated, almost all (scientific) char data I have contains non-ASCII characters, especially greek letters and such. Same goes even for web development or practically any user-oriented job where you need to accomodate for an international bunch of users. If you have pure ASCII data, lucky you, but why bother with UTF-anything then??
Also, the text I process is never UTF-16 at the source. So even if I used UTF-16, I would have to convert the text first and that would be the one and only time characters are counted. There would be no additional counting overhead at all.
Also, if that mattered to me, I'd store the char count in addition to the string length.
"UTF-16 is better": but 'only' for Asian languages; most code is still written in ASCII
"need to commit to 16-bit char width": the OP's point was that many many tools don't operate on 16-bit char width and with UTF-8 you can continue to use them.
The other two points are no advantages for UTF-16.
The authors of Go and authors of UTF-8 are more or less the same people, so the choice was no-brainer.
Edit: Ah, guys it's actually UTF-16, the configure flag is just named ucs2. False alarm.
//edit: this is about Java
Edit: Also, turns out it's UTF-16. The configure flag is named ucs2.
You are right though (and this is why I upvoted you back to 1) that you shouldn't care. In fact, you not knowing the internal encoding the proof of that. In python (I'm talking python 3 here which has done this right), you don't care how a string is stored internally.
The only place where you care about this is when your strings interact with the outside world (i/o). Then your strings need to be converted into bytes and thus the internal representation must be encoded using some kind of encoding.
This is what the .decode and .encode methods are used for.
Have a look at http://diveintopython3.org/strings.html which manages to say this better (and with more words) than I ever would be able to.
In practice the characters that aren't in UCS-2 tend to be characters that don't exist in modern languages, e.g. the characterset for Linear B, Domino tiles, and Cuneiform, so they're not supported since they're not of practical use to most people. There's a fairly good list at http://en.wikipedia.org/wiki/Plane_(Unicode) . In this list, Python by default doesn't support things not in the BMP.
Meaning: Every UCS-2 document is also an UTF-16 document, but not the reverse (just like every ASCII document is also an UTF-8 document).
But as I said below: It doesn't matter and could even be a totally proprietary character set as long as pythons string operations work on that character set and as long as there's a way to decode input data into that set and encode output data from that set.
>>> unichr(0x10000)
------------------------------------------------------------
Traceback (most recent call last):
File "<ipython console>", line 1, in <module>
ValueError: unichr() arg not in range(0x10000) (narrow Python build)
If you want to support codepoints greater than 0x10000 you have to recompile with the option UTF32.I think it must be a constant-lenght encoding to allow s[i] to be constant time.
The strange thing is that I couldn't find any reference to surrogate pairs in the Python documentation, so I was assuming that the elements of an unicode strings were complete codepoints. Instead this is not the case:
>>> list(u'\U00010000')
[u'\ud800', u'\udc00']
If I had Python compiled with the UTF32 option, this would return a single element, so Python is leaking an implementation detail that can change across builds. That's really really bad... u'\U00010000'.encode('utf-8')
should produce the same result on every Python version.Why? I am using list only to show what are the values of s[0] and s[1].
What I am saying is that it returns the list of characters of the underlying representation, so a list of wide chars (possibly surrogate) if compiled with UTF16 or a list of 32bit characters if compiled with UTF16.
Are you suggesting that all the string processing (including iteration) should be done on a str encoded in UTF8 instead of using the native unicode type?
A bit broader: too few programmers understand the difference between a character set and a character encoding.
why then they have the same names??? :)))
ps/ http://www.grauw.nl/blog/entry/254 - is this article ok? first in google by "charset encoding difference"
A character encoding is a mapping/function/algorithm/set of rules which can be used to convert a string into a sequence of bytes and back again.
A character set may have multiple encodings. UTF-8 and UTF-16 are two possible encodings of the Unicode character set.
The fact that the HTTP RFC speaks of 'charset=utf-8' is explained by this part of the spec:
Note: This use of the term "character set" is more commonly
referred to as a "character encoding." However, since HTTP and
MIME share the same registry, it is important that the terminology also be shared.
Why does MIME use the 'wrong' terminology? Perhaps because the registry is old and the difference between set and encoding was less obvious and relevant back then. Perhaps it was simply a mistake; a detail meant to be corrected. Perhaps the person that drew it up was inept. Who knows. It doesn't matter, it is still wrong. And don't get me started on the use of character set in MySql...// upvoted all replies, you're right
1. Most UTF-8 string operations can operate on a byte at a time. You just have realize that functions like strlen will be telling you byte length instead of character length, and this usually doesn't even matter. (It's still important to know.)
2. UTF-16 is still a variable-width encoding. It was originally intended to be fixed-width, but then the Unicode character set grew too large to be represented in 16 bits.
Probably because they figured they could just ignore endianness issues and that ASCII compatibility would be Somebody Else's Problem.
There were always problems with UCS-2. UTF-8 would have had a number of advantages over it even if Unicode had never grown beyond the BMP (Basic Multilingual Plane, the first and lowest-numbered 16-bit code space).
what are the issues?
> ASCII compatibility would be Somebody Else's Problem
for many of those outside "A" in ASCII (euphemism for America :) there were already a ton of problems, so endianness was the least (i personally never hit this problem)
// disclaimer: i'm not that serious about predominance of Latin script, this is sorta irony
It's easy to dismiss if you have all the time in the world and a deep stack of abstractions.
If you're doing deep packet analysis on UTF-16 text in a router, things may be different.
i'm not a native english speaker and a newb to HN, so sorry that i put my sincere question so that it looked like arrogant statement 'there are no issues, what are you talking about, i even don't know what LE and BE mean'.
will learn.
Abbreviation for 'American', in fact. No euphemisms needed.
(ASCII = American Standard Code for Information Interchange)
> there were already a ton of problems, so endianness was the least
I can appreciate this. However, UTF-8 also has desirable properties like 'dropping a single byte only means you lose one character, as opposed to potentially losing the whole file', and 'you can often tell if a multi-byte UTF-8 sequence has been corrupted without doing complex analysis'.
> i'm not that serious about predominance of Latin script, this is sorta irony
Heh. ASCII can't even encode the entirety of the Latin script: Ask a Frenchman how he spells 'café', or a German how he spells 'straße', and notice how important characters are missing from ASCII.
As a consequence at least Gecko and Webkit both use UTF-16 for their string classes, though there has been talk of trying to switch Gecko over to UTF-8. The problem then would be implementing the JS string APIs on top of UTF-8 strings efficiently.
http://download.oracle.com/javase/6/docs/api/java/text/Norma...
update: for particular purposes consider using Collator class, it makes collation keys (byte arrays) out of strings applying locale, case sensitiveness and unicode decomposition. (at least so says the doc, http://download.oracle.com/javase/6/docs/api/java/text/Colla... )
One being that's fair for all languages with respect to size so when you may be storing your standard Chinese, Korean, Japanese characters.
When UTF-16 made UC2 variable length, a few of the nice things were lost, but when dealing a lot of the higher code point characters mostly, UTF-16 may save you space.
utf-8 is good for network interchange and is de facto becoming standard.
utf-16 is not bad for internal storage of strings in memory or in database. not nesserily bad. maybe even better for some reasons
Edit: Your question might be answered at SO: http://stackoverflow.com/questions/1838170/what-is-internal-...
For external IO, internally it's in UTF-16 (actually UCS-2, mostly unaware of surrogates) or in UCS-4 (via a compile-time switch).
See http://docs.python.org/c-api/unicode.html#Py_UNICODE for the Py_UNICODE API.
[1] "Latin characters" is the proper term, not "American characters"
anyway, i think that even if Latin chars weren't the most used in the world, it would be fair to keep them the primary charset for use in programming and markup languages, as no-one now complains that the international language of medicine is Latin, not, say, Chinese :) as computers started to be massively developed in America.