Plain Text Doesn’t Exist: Unicode and encodings demystified
10kloc.wordpress.com
10kloc.wordpress.com
It's about time we actually start pressing for the idea that utf-16 was a terrible idea and that utf-8 should be the dominant wire format for unicode, with ucs4 if you really need to have a linear representation.
Utf-16 is confusing, complicated, and implementations are routinely broken because the corner cases are rarer. I really hope we're not still stuck with it in 50 years.
For example, consider a function to open a file by name. If your strings are UTF-8, you can just pass your null terminated buffer to fopen() or something and things will work fine for most files. But if your strings are internally UTF-16, you have to think about which encoding to use, and you research it, and you discover that holy crap, this stuff differs across OSes, and so we better take this problem seriously.
I'm really curious which languages drove you to this conclusion. Right now we're in a state where the most mature implementations of unicode in general are utf-16 because of the ucs2 accident, so it wouldn't exactly surprise me, but I'd still like to see the proof.
But it's not like utf-16 was invented to solve that problem so much as it was an accident that in the early days of unicode 2 bytes was enough to encode all code points and a 2 byte encoding was appealing as a universal solution in a way that a 3 or 4 byte encoding never would be. Some group of people on the planet will have to suffer 3 or more byte encodings no matter what. And I'd be happy to make it mine if it meant we could go with the better technical solution, if I had the power to do that.
All that aside it's certainly not quite as clear cut as any language. Many languages have a mix of ascii and non-ascii characters and those fare better than they would with utf-16. And in the codepoints below 0x07ff (where utf-8 goes up to 3 bytes) there are character sets like greek, hebrew, arabic, and cyrillic. And as has been pointed out, ascii characters are often used in structural elements of documents and this factors into making it unlikely that the general case is as bad as the worst case (50% inflation) for most practical text. Particularly when you throw compression into the mix.
[edit] fixed some minor things.
Looking at the format definition[1] anything below U+0800 will fit in 2 bytes in UTF8. So the ethnocentrism starts at Samaritan[2]. The Japanese, Chinese and Indian scripts are probably the most important in that range.
The comment you're replying to says right there, "implementations are routinely broken because the corner cases are rarer" utf-16 misleads people into using fixed-width code where it won't work right.
>Applications that use Asian languages take a big hit from utf-8.
Not really. It's a moderate hit for pure text, but how often do you see bulk pure text, without any markup, and that can't be compressed?
This is exactly my point: lots of code also treats utf-8 as if it were ASCII.
> how often do you see bulk pure text, without any markup, and that can't be compressed?
That's a good argument. Thanks :)
Two things to this:
- In many cases this is an entirely non-destructive mistake to make. This is an advantage of utf-8, that intermediaries that only care about ascii characters (even in the presence of extended characters) can work with it.
- It will become readily apparent that you've got a problem on the first attempt to port your app to any language with a non-ascii script. With UTF-16 it only becomes apparent when you go beyond the BMP.
Not really a myth. This was UCS2 and was the situation when a load of important early adopters started with Unicode. Winodws, Java, JavaScript all got burnt by this and ended up with UTF-16 as a result. Even Python 2.x on Linux is UTF-16 under the covers :(
>Unicode is just a standard way to map characters to magic numbers and there is no limit on the number of characters it can represent.
Unicode now limits itself to 21 bits of data. This is what allows the surrogate pair coding of UTF-16
There is plain text, and it is utf8.
http://web.archive.org/web/20090627072117/http://www.jbrowse...
[1] http://blogs.msdn.com/b/michkap/archive/tags/unicode+lame+li...
http://www.joelonsoftware.com/articles/Unicode.html
Though not exactly plagiarised, it shares some obvious parallels.
0 to 127. 127 is a power of 2 minus 1, which should be a hint; in specific, it's two to the seventh minus one, since ASCII defines codepoints for all possible combinations of seven bits, which is 128 possible codepoints, so the enumeration ends at 127 if you count starting from zero, as computer programmers are wont to do.
> all possible 127 ASCII characters
128 characters, as mentioned above.
> the ASCII guys, who by the way, were American
ASCII stands for American Standard Code for Information Interchange. The ethnocentrism was unfortunate but it isn't like you weren't warned.
> The numbers are called “magic numbers” and they begin with U+.
He can call them "magic numbers" but everyone else calls them "codepoints".
> UTF-8 was an amazing concept: it single handedly and brilliantly handled backward ASCII compatibility making sure that Unicode is adopted by masses. Whoever came up with it must at least receive the Nobel Peace Prize.
I'm sure Ken Thompson and Rob Pike will be happy to hear someone thinks that way.
To expand on this a tiny bit, a power of two minus one is called a "Mersenne number". If the number is prime, it's called ... wait for it ... a Mersenne prime.
When expressed in binary, Mersenne numbers are an uninterrupted series of one digits: 11111... of varying lengths.
> I'm sure Ken Thompson and Rob Pike will be happy to hear someone thinks that way.
If that's the case then I think the designers of GB18030 deserve it more, because they achieved an encoding that is able to map all Unicode codepoints while being backwards compatible with GB2312, which is itself backwards compatible with ASCII.
But seriously UTF-8 is like sliced bread after having dealt with we-thought-64K-is-enough-so-lets-all-use-16-bits UCS-2/UTF-16.
Of course, there's a long ways to go from "a sequence of valid code units" to "a valid string," but it's still relevant.
what about surrogate pairs? You can't have only one 16 bit word for a pair and have a valid UTF-16 sequence. This problem is real easy to do if you substring a UTF-16 sequence naively.
In what sort of situation that is a problem?
There's a related issue with "non-shortest forms:" it's possible to encode the NUL character '\0' without actually using a null byte. This means that (for example) "round tripping" UTF-8 data can introduce a null byte, allowing for security exploits when mixed with C-style string processing.
But what I wrote is nevertheless correct, and matters. See for example the special requirements on UTF-8 processing in D93:
If the converter encounters an ill-formed UTF-8
code unit sequence...it must not consume the
successor bytes as part of the ill-formed
subsequence whenever those successor bytes
themselves constitute part of a well-formed UTF-8
code unit subsequence....For a UTF-8 conversion
process to consume valid successor bytes is not
only non-conformant, but also leaves the converter
open to security exploits."
"Ill-formed code unit sequence" is not possible with UTF-16. That's an advantage. Also note the last line: the existence of ill-formed code units has security implications that are not present with UTF-16.To be fair, UTF-16 has its own unique issues too, like endianness.
Yea, I think it reflects how much technology development comes from America even back then as well as now.
https://plus.google.com/101960720994009339267/posts/Rz1udTvt...