Emoji was added because Apple and Google had to deal with (then-)Japanese emails, and skin tones were not specified. It were implementations that impose certain skin tones (that do not even match the original Japanese emojis) and as a result Unicode had to introduce a mechanism to change skin tones and mandate the default emoji without that mechanism to be neutral.
RTL "modifiers" are actually formatting characters closely tied with the Unicode Bidirectional Algorithm [1]. Until then texts with both RTL and LTR fragments were handled incoherently, for example legacy character sets were still struggling with logical vs. visual order issues. So they are indeed Unicode inventions, but necessary ones that do not alter existing texts.
For CJK characters Unicode now provides ideographic variation selectors that select the exact glyph (or more accurately, a restricted glyphic subset of the base character). They do not disunify characters but they do provide a strong hint to display those characters in a specified way. In this way they do not cause an additional issue to existing Unicode systems (as they should already do normalization and collation in the Unicode way). The disunification by comparison would almost instantly break existing texts.
Adding skin tones was a choice, there was no immediate need for it.
You are incorrect. See the original design document [1] for skin tones and other diversity improvements, especially the "Sources of input" section.
[1] https://www.unicode.org/L2/L2014/14172r-emoji-enhancements.p...
So what do we do with 国 and 國? The first of those is always used in simplified Chinese and usually in Japanese, while the second is used in traditional Chinese and sometimes in Japanese (eg. names). Is this one, two or three characters?
More on the topic: https://en.wikipedia.org/wiki/Han_unification
I'm personally of the belief that the accented and cedilla characters should be exclusively stored as combining character pairs, even if modern keyboard mappings require only a single keypress. My own language stores every character as two bytes (at a minimum), so the storage aspect is a solved problem.
If all three had different codepoints and you replaced 国 with 國 a lot of people would realize, less so if you replaced 國 with 國.
To my understanding the only argument in favour for han unification was that it would have taken-up a lot of codepoints otherwise.
I can assure you that Han unification happened at the hand of native speakers.
Did you know that some languages distinguish dot less Iı and dotted İi ? English mixes them Ii and unicode needs to know exactly based on what language you might want to upper/lower case because it can't tell an english I appart from a Turkish dotless I.
The CJK variant issues under not specifying a lang are indeed present in Latin but to a smaller extent.