> I do not like UTF-8 / 16. It's effectively bad huffman encoding. It's an attempt to save space, but it doesn't even do that well.
UTF-8 + gzip is 32% smaller than UTF-32 + gzip using the HN frontpage as corpus. Even using xz, it's a 13% gain.
> About the only advantage of UTF-8 is that ASCII maps to it reasonably well.
That's a pretty huge advantage, and a big reason why UTF-8 is actually popular. An other one is UTF-8 being byte-based, it does not care for byte order. UTF-32 is split between BE and LE, and requires either out-of-band byte-order communication or a BOM.
> Why on earth does any higher-level language still use byte or codepoint counts for length?
Because it's easy, and generally O(1) in these languages. Can also be useful to know how much space it'll take when stored, which really is the only useful use for a string length.
> And why don't lower-level languages at least have a way to count / index by graphimes?
Counting graphemes is no more useful than counting bytes or codepoints. You could provide a grapheme cluster count, but:
1. that's O(n) period
2. it serves very little purpose since clusters don't have a fixed width, not even with a fixed-width font
3. clusters can be locale-dependent ("tailored" clusters) although the default set is locale-independent. Now you need to ponder whether you include tailored clusters, don't include them, or optionally include them
4. clusters and glyphs are independent, "ch" is a grapheme cluster in Slovak but two glyphs on-screen, whereas an "fi" ligature is a single glyph but two clusters