Languages that use UTF-16 mostly use it only because they're older than UTF-8 (or are tied to a platform older than UTF-8).
I think you're right that older languages use UTF-16 and newer ones use UTF-8. But it also seems empirically true that UTF-16 languages do better at grappling with Unicode's subtleties, compared to UTF-8. The temptation of UTF8-is-bsically-C-strings is hard to ignore.
You have to deal with unpaired surrogates, though. And just because that wasn't annoying enough, any UTF-8 sequence which contains an encoded surrogate (which is technically invalid, but not prohibited by most implementations) is impossible to encode as UTF-16.
What Javascript does is use UTF-16 semantics for Unicode strings. The reason why it does this is simple: when those methods were implemented in the mid-1990s, UTF-16 was largely synonymous with Unicode. No characters beyond U+FFFF were defined until the release of Unicode 3.1 in March 2001.
With UTF-8 you'll at least have a shot at noticing that you're not handling multi-unit codepoints well, while with UTF-16 you won't notice unless you test Chinese or a more off the beaten path language.
Regarding your last point, I'm totally with you (if I understood you correctly). Of course applications should support multiple input/output encodings, but as a programmer you have to decide on some internal representation.
That being said, I really don't see how processing UTF-8 is significantly more complex than processing, say, UTF-16. In both cases you need to handle continuation units for the extraction of Unicode code points.
UTF-8 has 4 valid cases, one for each length, and many more invalid cases for each length (2-byte sequence missing trail byte, 3-byte sequence missing 1 trail byte, 3-byte sequence missing 2 trail bytes, 4-byte sequence missing 3 trail bytes, 4-byte sequence missing 2 trail bytes, ..., overlongs, UTF-8'd surrogates, overflow, etc.) Differences between implementations' treatment of error cases have lead to some security concerns; see https://hsivonen.fi/broken-utf-8/ and discussion at https://news.ycombinator.com/item?id=14451822 for an example.
UTF-16 has two valid cases (one or two code units) and two error cases (lead surrogate not followed by trail surrogate, lone trail surrogate). It's more like a DBCS, except each code unit is 2 instead of 1 byte.
As someone who has actually written UTF-8/UTF-16 conversion code, I can immediately tell you which one is far easier to implement: UTF-16. The number of valid cases is basically halved, and the number of error cases in UTF-16 is a fraction of those in UTF-8. Put another way, there are plenty more invalid UTF-8 sequences than invalid UTF-16 sequences.
In any case, the discussion here is the appropriate string API, and the relative difficulty of working with those. Exposing UTF-8 versus UTF-16 changes essentially nothing: in both cases you need to either deal with non-integer indexes or deal with integer indexes where not all values are valid.
Good string APIs are hard. Most Unicode-aware languages pick one particular encoding and then toss the programmer in the deep end with it.
The only language I've seen get it vaguely correct is Swift. (I'm sure there are others, but it's definitely not common.) Swift strings provide multiple views, so you can work with UTF-8, UTF-16, UTF-32, or grapheme clusters, as you need. It doesn't allow using integer indexes directly, so you have to confront the fact that indexing is actually non-trivial. Swift 3 requires using views, and Swift 4 makes the String type itself a sequence of grapheme clusters, which is usually the correct answer to the question of "what unit do you want to work with?"
In my experience, I have not yet found a case where I ever wanted to use grapheme clusters. Most algorithms want to iterate over Unicode codepoints (e.g., displaying fonts). Even in display cases, grapheme clusters isn't necessarily the right thing to use for the backspace key or left/right motion.
Plus I'm pretty sure all the major JavaScript engines (V8 for sure) already know how to handle UTF-8 since that's the encoding most scripts come from.
UTF-16's validation concerns are:
1. Broken surrogate pairs, which is mostly benign.
2. Byte-order confusion.
While UTF-8 has:
1. Invalid code points, for example, code points for surrogate halves.
2. Invalid code units, such as 0xFF.
3. Non-shortest forms, where a character may be encoded multiple ways.
4. Representation of NUL, and potential for confusion with APIs that expect null-terminated strings.
In practice the UTF-8 issues have caused much more serious vulnerabilities.
In any case, these are all concerns for a decoder, but not for an API, which is what we're discussing here. In fact, the original comment I replied to up there was advocating the opposite: UTF-16 internally, and UTF-8 for interchange!
while (*c) count += ((*(c++) & 0xC0) == 0x80) ? 0 : 1;
See https://stackoverflow.com/questions/9356169/utf-8-continuati... for more details.Counting the number of UTF-8 code units in a UTF-8 string is of course trivial. Counting the number of UTF-16 code units in a UTF-8 strings would take more work. But there's probably no reason you'd want to compute that anyway.
Back then, the trade-off made sense, and was popular too, because variable-sized encodings were much more rare. These days UTF-16 really is the worst of both worlds, but we are stuck with what we have.