> and Japanese is basically a hybrid of those two.
This is prime /r/BadLinguistics fodder but you've ironically hit the head on the problem.
The underlying issue is that Unicode was run by people who thought 16 bits was enough[0], ran into the issue of Chinese characters, and imposed a bunch of very specific unification rules to work around their own self-imposed technical limitation so they could retain 16-bit codepoints. The rule is that characters that only differ in appearance are treated as the same character[1].
To explain how dumb this is, I'm going to invent a concept called UniPhoenician. You see, Latin, Greek, Cyrillic, and a few other phonetic scripts have significant derivation from Phoenician writing, so we're going to just merge the letters that happen to have a shared pedigree. i.e. Latin a, Greek alpha, and Cyrillic a. Of course, once we do that, software now has to be consciously aware of what language text is written in so it can substitute the right set of glyphs to work around UniPho.
To make this even dumber, the limitation that motivated UniHan went away with Unicode 2.0, which went to 20-bit codepoints. Except we didn't fix UniHan, AND we subtly broke old software. You see, 16-bit codepoints was the only encoding for Unicode 1.0. Unicode 2.0 added UTF-8[2] and UTF-16, the latter of which is a series of rules on how to fit 20-bit codepoints into 16-bit text in a way that subtly breaks old software, which hopefully will get updated and then people can just pick what codepoint length they want.
Well, uh... turns out Windows NT and JavaScript already were using 16-bit codepoints, and integrating the new UTF-16 rules into them subtly breaks existing software based on that. So those can never be fixed, and any software built on their text-handling capabilities is subtly broken in the face of emoji, rare characters, and so on. Bonus points is that, because they can't understand UTF-16's special characters, naively written conversion functions working with 16-bit-only software will leak invalid UTF-16 sequences into UTF-8, as documented in WTF-8[3].
[0] Competing proposals for a universal character set, including ISO's UCS, used 4 byte characters, see:
https://en.wikipedia.org/wiki/Universal_Coded_Character_Set#...
[1] Unless the difference is between Simplified and Traditional Chinese, because Unicode didn't wanna piss off Mainland China but was okay with pissing off Japan and Korea
[2] AKA Filesystem Safe Unicode, which shoves Unicode in 8 bits in a way that is mostly acceptable and doesn't impose any Latin-centrism that non-Latin script users need to worry about. Bonus points is that it was originally designed for 32-bit codepoints (anticipating UCS?). If we ever needed 32-bits, UTF-8 could handle them, while UTF-16 would require additional layers of hacks that would spill over into WTF-8.
[3] https://simonsapin.github.io/wtf-8/