So then we all end up paying this massive complexity tax everywhere to pay for support for some Mongolian script that died out 200 years ago (or multi codepoint encodings of simple things like é - just why, it was so avoidable).
So then we all end up paying this massive complexity tax everywhere to pay for support for some Mongolian script that died out 200 years ago (or multi codepoint encodings of simple things like é - just why, it was so avoidable).
This is not true. For a concrete example: the languages Hindi and Marathi, with ~500 million speakers, use the Devanagari script (also used by Nepali and Sanskrit), in which a grapheme cluster is (usually) a sequence of consonants followed by a vowel. For instance, something like "bhuktvā" (भुक्त्वा) would be two grapheme clusters, one (भु) for "bhu" and one (क्त्वा) for "ktvā". In Unicode each vowel and consonant (here, bh, u, k, t, v, ā) is separately encoded, which is the only reasonable thing to do, and inevitably means that grapheme clusters can have different lengths (number of code points). The alternative would have been to encode every possible (sequence of consonants + vowel) as a single codepoint, which gets ridiculous quickly: these sequences can be up to 5 consonants long, so you'd end up having to encode (33^5 * 13 ≈ 500M) codepoints for Devanagari alone (or completely prevent certain sequences of consonants from being expressed, which makes no sense either), not to mention that most of the scripts of the Indian subcontinent and south-east Asia follow the same principle and have similar issues (e.g. Bengali with 250M speakers, Telugu, Javanese, Punjabi, Kannada, Gujarati, Thai with over 50M speakers each, etc).
(See chapters 12–17 of the Unicode standard, currently version 15: https://www.unicode.org/versions/Unicode15.0.0/ch12.pdf)
It would be nice if we could come up with some magical system that optimally encodes all the text that "matters" and ignores everything else, but history has shown that to be very hard. So we're left with Unicode, which takes the approach of giving us (effectively) infinite code points to represent characters, with (effectively) infinite ways to visually represent them. That does lead to a bunch of "unnecessary" baggage and headaches, but it also solves a bunch of real problems that you probably don't know exist.
Unicode is a pain in the ass, but it's a solution to a very hard problem. You can feel free to design your own solution, but you'll probably run head-first into all the problems Unicode was trying to solve from 40 years ago.
That said, what it's trying to do is enormously complex.
P.S.: Also, even for those, it would seem that one of the big reasons for things like combining characters was added to Unicode in order to be backwards compatible even with mutually incompatible encodings ?