Most server-side applications should never have to know what these concepts even are. Or any library that is not user-facing. They get bytes from the UI layer, and they can keep them as opaque bytes. For user-facing apps, you can ask your renderer library for a pixel-width or similar for a string, and let them handle how to parse it. Very little code ever needs to know about unicode.
Any kind of input-sanitization is vastly simplified in utf8, and that makes it worth it for me. For me the really troubling trends are conventions like Rust Utf8Error, where they can cause what I'd consider a UI-related exception in code that had no business even interpreting what those bytes are. Unfortunately, every API uses strings, so they are kind of hard to avoid. It introduces what I'd consider a software layering problem.
Maybe others here with more experience with internationalization can chime in and tell me I'm wrong.
> For me the really troubling trends are conventions like Rust Utf8Error,[…]
Interesting. Isn't that only returned when the input bytes contain a non-UTF-8 byte sequence? How is Rust's approach different from other languages?
As a Rust user: this is what your function returns if the user inputs non-UTF-8 bytes into something that expects UTF-8 bytes and the programmer explicitly choses not to handle that error.
I don’t see anything wrong with that. Sometimes you might be interested in receiving, processing, storing valid UTF-8 strings rather than arbitrary byte sequences, that may or may not be able to be translated back into something valid that you can display.
I always hated to do encoding related work with a passion before I started using Rust. Rust forced me to do it the right way and actually understand why I am doing it a certain way. When it comes to encoding I feel safer in Rust then e.g. in Python, despite having used it for four times as long.
This is a common misconception with UTF encodings.
> A character doesn't mean anything in Asian languages, and any attempt to use a fixed-length encoding is pointless.
Totally agree, indexing into a string is bad practice, no matter the encoding (because you probably can't ever guarantee what the encoding will be, or how it was converted, etc.) This is true for UTF-8 and UTF-16, and really any encoding because again, you can't be sure what you're dealing with. Code points, "characters", glyphs, etc. are all concepts that work at different parts of the stack (well, characters doesn't), which again is true for all encodings.
The advantage UTF-16 has for languages that tend to be multibyte is it's representation takes up less space. Other than that, it has all the same disadvantages any other encoding has.
> Any kind of input-sanitization is vastly simplified in utf8, and that makes it worth it for me. For me the really troubling trends are conventions like Rust Utf8Error, where they can cause what I'd consider a UI-related exception in code that had no business even interpreting what those bytes are. Unfortunately, every API uses strings, so they are kind of hard to avoid. It introduces what I'd consider a software layering problem.
This is the problem right here. The "UTF-8" the world initiative ignores that UTF-16 is a lot more practical for most people, and as a result you get platforms expecting UTF-8 that really have no business doing so.
It's a partial fix that makes bugs harder to catch. There's also the BOM issue, and not being ASCII compatible...
Consider things like databases (relational/document/key-value/etc.), JSON payloads, compression algorithms, battery life, etc. etc.
There is a problem with lots of developers assuming UTF-16 is fixed-width (just look at this thread), but that's a mistake for any UTF encoding, and it's almost always a mistake in any encoding--you shouldn't be indexing into strings.
People say, "just use UTF-8, you never have to worry about it and it's ASCII compatible", but you shouldn't be using the ASCII compatibility (arguably this is an anti-feature) and you shouldn't have to worry about the details like BOM because you're using your language's built-in string support or a library right?
Where there is significant markup, UTF-16 encoding is rarely more space-efficient than UTF-8. When compression is involved, there is rarely a difference.
ASCII: 1 byte
European characters: 2 bytes
Asian characters: 3 bytes
There is no reason you'd want this in Asia. China uses GB2312, which flips it around: ASCII: 1 byte
Asian characters: 2 bytes
European characters: 3 bytesNote that “Asian” is quite a gamut in terms of UTF-8 size: Chinese is at the very compact end of the spectrum and the Burmese script at the other end: https://hsivonen.fi/string-length/#counts
I'm a little skeptical about that table though, because as you point out, there's not a great way of figuring "meaning per character". For some more (real flawed) comparison, the English version of The Tale of Genji is ~60k words and 224 pages, and the Chinese version is ~75k words and 300 pages. [1] [2]
So yeah, I'm skeptical about the claim that some languages have more meaning per character, and as a result you'll end up storing less text overall. I think it'd be cool to look at more data about this though, like (for example) stats from Treasure Data [3].
But regardless, it's pretty indisputable that UTF-16 is a lot better space-wise for languages that tend to be multibyte in UTF-8. Mostly what I'm trying (Quixotically) to say is "UTF-8 the world" ignores a lot of the world.
[1]: https://www.readinglength.com/book/isbn-4805314648
Note that link [2] says "guess [of # of words] based on page count". 75k words, 300 pages is one statistic, not two separate statistics. (Estimating words based on page count is likely to be highly reliable, but still.)
> I'm skeptical about the claim that some languages have more meaning per character, and as a result you'll end up storing less text overall.
Your skepticism is unwarranted. From a pure information-theoretic perspective, the claim that some languages have more meaning per character is a slam dunk, and less than a second of examination proves it conclusively.
For example, an entire English novel is unlikely to use more than 256 unique characters. A Chinese novel couldn't use anywhere near that few without sounding incredibly artificial.
Here are some single-character words in modern Mandarin:
高 - high/tall
低 - low
大 - big
小 - small
最 - most (superlative marker)
更 - more (comparative marker)
到 - arrive
走 - leave (go away)
玩 - play (e.g. a game)
看 - look
听 - listen
贵 - expensive
爱 - love (verb)
恨 - hate (verb)
龍 - dragon
For something more representative, here are some song lyrics -- I'll enclose every word of multiple characters. Unenclosed stretches of text are words of one character each: 都怪我不[小心]地
[偷看了]你的[眼睛]
[迷失了][自己]
[不知][为什么][欢喜]
[期待]听你的[声音]
是[那么][甜蜜]
我[拿着][电话]在你家门外
想对你[说出]我[内心]中对你的[期待]
[忽然][感觉][浑身][热血][澎湃]
[今天]我[想要][[鼓起][勇气]][大声][表白]
How many distinct one-character words do you think English could practically support? How many twos? In this verse-and-a-half, there's one word that's always[1] three characters and two that have reached three characters by picking up a verb suffix, for a total of three words that are longer than two characters. (鼓起勇气 is kind of a special case, in that it's a well-known fixed expression, but its meaning is transparent as an ordinary combination of the two words 鼓起 and 勇气, which also see use outside the expression.)[1] Actually, 什么 ["what"] has the vernacular contraction 啥, and this also applies to 为什么 ["why"], so you could argue that the word is sometimes just two characters.
Yeah I mean, there's no question ideographic languages have more meaning per character than alphabetic languages. But other comparisons and considerations aren't as obvious:
- how do ideographic languages compare with each other?
- do people using ideographic languages write more?
- are ideographic languages as effective at compression as general (or special) compression algorithms?
Moving up the conceptual ladder from "average bytes per codepoint" to "average size of encoded tweet" (for example) is a big leap is all I'm saying.
That one we know; the compression algorithms are more compressive. For example, compressed Chinese text takes up less space than the same text uncompressed. Ideographic languages are still languages that real humans have to use, and they feature redundancy because that helps everyone. Compression algorithms have the luxury of stripping that redundancy out.
> how do ideographic languages compare with each other?
Of modern languages, only Chinese and Japanese could really be described as ideographic. (Japanese much more so than Chinese, in fact.) Chinese will have more meaning per character, because Japanese makes heavy use of the comparatively less meaningful kana. (Interestingly... it has to do this precisely because of its more ideographic nature.)
> do people using ideographic languages write more?
No idea.
A related question is if people using ideographic languages write electronically when they do, and I suspect yes.
Language Log has several articles on the phenomenon where Chinese seem to be forgetting how to write characters due to IT input methods. https://languagelog.ldc.upenn.edu/nll/?p=7142
Handwriting is actually a moderate-level problem for me -- I can't recognize most handwritten characters. Often a special handwriting form will be used.
OK, I lied. China doesn't use GB2312 anymore. They use its update, GB18030, which carries arbitrary Unicode data just as UTF-8 does... except that, like GB2312, it puts the Asian characters in two bytes and the European characters in three.
It's certainly not obvious to me that UTF-8 is better for globally appropriate software. It looks worse. Software for Europeans, sure.
There are more important qualities than optimizing byte length. Processing GB18030 as Unicode scalar values involves lookup tables. A single byte error cascades potentially further than in UTF-8.
If optimizing Chinese byte length is really important for you, UTF-16 is easier to work with than GB18030.
> When compression is involved, there is rarely a difference.
Rare is the situation where smaller input to a compression algorithm leads to larger output, and UTF-16 is smaller in lots of languages. It might rarely make a difference you, but that's not the same as rarely making a difference.