I would like to understand why UTF-32 didn't catch on as The Standard Unicode for the modern world. it seems that - albeit memory wasteful - it would sidestep a lot of these issues.
I would like to understand why UTF-32 didn't catch on as The Standard Unicode for the modern world. it seems that - albeit memory wasteful - it would sidestep a lot of these issues.
> memory wasteful
The answer is in the question really. If you've got a big pile of mostly-ascii data, quadrupling memory/storage to encode it as UTF32 is going to be a pretty tough sell
(If you’re doing number-crunching on giant CSVs, maybe I can see it being important, but all the ascii files on my desktop that I can think of are pretty trivial)
> UTF-8 + gzip is 32% smaller than UTF-32 + gzip using the HN frontpage as corpus.
My use case was to show the flaw in the logic a colleague was using to assert that gzipped JSON should be the same size as gzipped MessagePack for the same data, because the information content was the same. It was a quick 5-minute script without having to deal with coming up with a suitable JSON corpus to convert to MessagePack.
Among other things, the zlib compression window only holds half as many characters if your characters are twice as big.
$ </usr/share/dict/words gzip --best | wc -c
261255
$ </usr/share/dict/words iconv -f utf-8 -t utf-16le | gzip --best | wc -c
303404
A bit over a 16% size increase form converting the "wamerican" dictionary to UTF-16LE and then compressing. $ < /usr/share/dict/american-english-insane gzip --best | wc --bytes
1778330
$ < /usr/share/dict/american-english-insane iconv -f utf-8 -t utf-16le | gzip --best | wc --bytes
2061457(UTF-8 is compatible with ASCII, but even if it was stored in some other encoding, conceptually it could be in ASCII, you know?)
Our world however runs on 8-bit bytes, so it makes some sense for text to be based on that.
But also, consider Base64 in UTF-32-encoded JSON. ;)
You just posted this to one.
https://lucumr.pocoo.org/2014/1/9/ucs-vs-utf8/
The favicon btw is cached and amortized across all HN pages whereas the text is not.
I forget where I read this but about a decade or so I remember reading a paper or watching a video (maybe from the Azul folks?) that looked into JVM memory usage and a good chunk of it was strings and the ucs2 encoding was a problem. That’s why even languages that are nominally utf16/32 as the native type will frequently auto detect and special cases latin1 strings (python, js, etc). The other piece of it is that strings are copied around more and processed differently from images. The knock on effects of utf32 can be quite unfortunate (ie rendering your html document is meaningfully slower which you care about even though by weight your images take longer to transfer and show)
I feel that this is worth a blog post in itself. I remember years ago comparing Go and C# when processing some large mostly-ascii files. The C# program was faster, much to my surprise, despite storing strings natively in UTF-16. (I don't remember the implementation details, so it may have been an artifact of my implementations).
Well, the use case mentioned in the article is a pretty good one: program source code. Even if you're going to be writing in a foreign language, all of the fancy punctuation and whitespace that does useful stuff in the language ends up being ASCII, and a good hunk of the standard library is likely to have ASCII names for types and functions, etc.
Most databases. It might be compressed on disk, built no DBA wants all their column lengths quadrupled.
Also it's important to look at the time period. The farther back in time you go the larger a percentage of all data was designed for direct human consumption. (This is why things like binary coded decimal existed over binary.)
They are not grapheme clusters, such as the "family: man, woman, boy" emoji from TFA.
¹which is approximately what I think you're saying here. I.e., you're trying to say that a code point might span multiple UTF-32 code units; that is not correct. (It should be simple to see how a code point, which has the range [0, 0x10FFFF], can always fit into a u32.)
Everything else has.
And then is UTF-16 which has all the pains of UTF-8 with none of the advantages of UTF-32
Officially, it's at most four bytes, of which 21 bits are usable for encoding codepoints - so that's an upper limit of 2^21 codepoints.
There is an initial byte encoding the length as a series of ones, so if you went ahead and extended the standard to simply allow more bytes, you could get up to 8 bytes, of which 48 bits would be usable.
I can see that a six-byte version with 31 data bits was previously standardised before they settled on four.
I guess you could extend it further by allowing more than one initial byte encoding the length, then it would be arbitrary length. But at that point I'm not sure if it loses its self-synchronising ability, and in any case it would be a different standard at that point.
I think you'd only be able to go up to 7, since 10xxxxxx is still reserved for trailing octets. And even with 7, the entire first octet is consumed by the length indicator alone.
So you get 0xxxxxxx, 110xxxxx, 1110xxxx, 11110xxx, 111110xx, 1111110x, and 11111110 as the 7 different length-indicating head octets. In the last case, you'd have 36 usable bits for encoding a codepoint.
Also note that if you did add 11111111 as a valid head octet representing an 8 octet long encoding, you'd still only have 42 usable bits (since the first byte is still entirely consumed by the length indicator)
More generally, every code unit in the file has to have the form 00xxyy00, and the possible values for xx and yy must be in the range 0 thru 16, so there are only 17*17-1 = 288 unicode code points that can possibly occur in a endianness-ambiguous string/file.
And out of those, you have confusions like "Ā" (U+0100 LATIN CAPITAL LETTER A WITH MACRON) versus "𐀀" (U+10000 LINEAR B SYLLABLE B008 A) or "𠀀" (U+20000) versus "Ȁ" (U+200), most of which can only happen if you have no idea what language you're expecting. (Some, like "𐄀" (U+10100 AEGEAN WORD SEPARATOR LINE), are the same regardless.)
Whereas here's a valid bash script:
⌀℀ ⼀戀椀渀⼀戀愀猀栀猀甀搀漀 椀搀⌀ 2>/dev/null || true
echo "Hello, World!" #
(You can almost do this with C/C++ as well, but something needs to #define "⼀⼀" as "int" for it to work properly, and then only if the compiler accepts non-ascii characters (namely "⼀" aka U+2F00 KANGXI RADICAL ONE) in identifiers.)People are surprised, confused, and sometimes even offended by the fact that I do almost all my work with a plain text editor and a terminal. I have the same sentiment towards those who insist upon large complex fragile stacks of tools and then wonder why they spend so much time chasing down bugs in those rather than working on what they actually intended to.
This is a bug in the less popular third-party lsp package for emacs, which is already quite unpopular.
I use VSCode, an enormously complex system. But so many other people use it there tends not to be this sort of bug. And in the rare case there is one I just wait a few days until someone else solves it.
Sure if, everyone used UTF-32 for everything then these problems would go away but they would also go away if everyone used UTF-8, and most uncompressed files would be 4 times smaller.