I'm fine with
characters having a variable number of bytes. The problem comes when I have to interpret individual characters and/or codepoints and/or glyphs in a UTF string when trying to correctly lay out text on a Linux Terminal.
I just find it disconcerting that a single character in UTF-8 can be of infinite length (an infinite chain of modifiers, followed by an emoji; or an infinite stack of composing accents). And that the Unicode consortium has chosen to break the one-code-point-per-glyph convention.
An example of a truly weird outlier: the "family" emoji, which is composed as ManEmoji+ZeroWidthJoiner+WomanEmoji+ZeroWidthJoiner+ChildEmoji, which produces a single glyph.
And then how to deal with modifiers? e.g.
MonochromeEmojiModifier+
AsianSkinToneModifier+
ManEmoji+
ZeroWidthJoiner+
WomanEmoji+
ZeroWidthJoiner+
ChildEmoji.
(70 UTF-8 characters, 14 UTF-16 characters, 7 UTF-32 characters! How many glyphs? From one to three depending on what typeface you're using!). Unlike composing accents, which have a combined code-point, the Family emoji does not have a combined code-point, so it is typically implemented using combining in glyphs in the output TrueType/OpenType font, since TT/OT fonts can combine glyphs into private codepoints visible only inside the font.
or
AsianSkinToneModifier+
ManEmoji+
ZeroWidthJoiner+
AsianSkinToneModifier+
WomanEmoji+
ZeroWidthJoiner+
AsianSkinToneModifier+
ChildEmoji.
(Is that even legal?)
And then an appendix of special rules that apply mostly to obscure scripts/locales that have to be dealt with almost on a case-by-case basis (fortunately, of those, German and Turkish are the only locales with odd special cases that Linux supports).
ALL of which are encoded in about 500,000 lines of code, documentation and data files by the Unicode consortium, that are updated annually. Unicode 18.0 includes 18 new emoji (including the infamous Pickle emoji, and skin-tone modifiers for left- and right-thumb), and three new scripts including proto-cuneiform. Yay! One assumes that Linux will implement the Pickle emoji on an urgent basis, and that proto-cuneiform will never be supported as a system locale.
The use-case that almost broke me: line-wrapping and displaying and editing arbitrary user-entered UTF-8 paragraph text on a Linux terminal -- both graphical (almost full Unicode support but very out-of-date emoji), AND true text mode (up to 512 locale-dependent glyphs, each of which may or may not be double-width) that has to be supported to allow use on machines without desktops. The IBM Unicode libraries get you part of the way there, but far from all of the way there; and the few system APIs that Linux do provide are buggy as heck.