What? No, UTF-8 won for a reason, and that reason is not just that C has a deficiency in this area but, rather, that UTF-8 is: a) simple and much saner than UTF-16, b) self-synchronizing in both directions, c) as -or even more- space efficient than UTF-16 on average even for non-Latin text, d) even UTF-32 doesn't make it possible to turn logical character string indices into 32-bit word string indices. (d) is the real killer.
One just cannot assume that a character or glyph requires just one codepoint to express, therefore one can't assume that a character or glyph will require some fixed number of code units to express, therefore one might as well use UTF-8 because it's saner and more efficient than the other UTFs, therefore... <drumroll/> a "constant clinging to and dependency upon the “multibyte” encoding" is NOT a deficiency of C but an advantage of C.
C does have problems here though, namely all the usual problems it has:
- C strings (NUL-terminated) suck
- C doesn't have a first class string type,
only pointer to char
- C doesn't have a first class string type
that indicates what encoding the string
uses
Now, C's wchar_t is not even a deficiency but a disaster.