> I just find it disconcerting that a single character in UTF-8 can be of infinite length (an infinite chain of modifiers, followed by an emoji; or an infinite stack of composing accents).
One of the Unicode guidelines suggests it's not very important to handle sequences that are more than 31 code points long.
> And that the Unicode consortium has chosen to break the one-code-point-per-glyph convention.
The what convention? Combining accent marks were in there from the beginning, and those came from previous text formats.
If you're dealing with a terminal you should be tracking one grapheme clusters per cell (or per two cells if you allow wide characters) and taking notice and control over when they merge. Store them in a way that putting line breaks in between the grapheme clusters is easy.
> And there are the truly weird outliers, like "family" emoji, which is composed as ManEmoji+ZeroWidthJoiner+WomanEmoji+ZeroWidthJoiner+ChildEmoji, which produces a single glyph. And then how to deal with modifiers?
That's the thing, it's not actually any weirder than the hangul jamo encoding for Korean text. And the flags also work like that but without needing ZWJ. They didn't add this complicated stuff for the sake of emojis, and like another comment said emoji acts as a shiny lure to get everyone to implement the spec properly.
> 70 UTF-8 characters
If you want to conflate characters and code points I'm not going to yelp too loud but please don't conflate characters and UTF-8 bytes, that's not how it works.
Also I want to know how you calculated that number because it's wrong. It's 28 bytes in all three of those encodings. Even a broken conversion from UTF-16 into UTF-8 that bloats the number of bytes is 42.
> And then an appendix of special rules that apply mostly to obscure scripts/locales that have to be dealt with almost on a case-by-case basis (fortunately, of those, German and Turkish are the only locales with odd special cases that Linux supports).
That sucks but humans made text hard long before computers were around.
> Unicode 18.0 includes 18 new emoji (including the infamous Pickle emoji, and skin-tone modifiers for left- and right-thumb), and three new scripts including proto-cuneiform. Yay! One assumes that Linux will implement the Pickle emoji on an urgent basis, and that proto-cuneiform will never be supported as a system locale.
That seems fine? Adding another emoji is really easy and adding a locale is hard and will only be done if there's users.
> The use-case that almost broke me: line-wrapping and displaying and editing arbitrary user-entered UTF-8 paragraph text on a Linux terminal -- both graphical (almost full Unicode support but very out-of-date emoji), AND true text mode (up to 512 locale-dependent glyphs, each of which may or may not be double-width) that has to be supported to allow use on machines without desktops. The IBM Unicode libraries get you part of the way there, but far from all of the way there; and the few system APIs that Linux do provide are buggy as heck.
It does suck that these libraries and APIs are a mess.