I found this old comment that explains it better than I can: https://news.ycombinator.com/item?id=28287811
I found this old comment that explains it better than I can: https://news.ycombinator.com/item?id=28287811
Unicode simply lists all possible combinations in dictionary order starting from U+AC00. So you can take any code point and split out the 초성, 중성 and 종성 using simple arithmetic, just like you can figure out Latin alphabets from their ASCII codes.
My understanding is that there are two possible unicode encodings of Korean, one of which (MacOS) is sound by sound instead of syllable by syllable (Windows). This is why Korean UTF-8 filenames from MacOS appear broken on modern Windows machines.
Having said that, MacOS also made the strange choice of expressing Hangul using the Hangul Jamo (by sound) Unicode block even when there are equivalent precomposed symbols in the Hangul Syllables block. Encoding each sound individually takes up 2-3 times more storage, just like with accented characters in Latin. Besides, if you just list sounds and rely on them to be combined automatically, what do you do when you legitimately want to write a sequence of uncombined sounds, like ㄱㅏㅁ instead of 감?
Sure, Unicode isn't the Platonic ideal of a character encoding. It has warts, legacy features, and.. and it is a universal encoding of all human writing. What an exceptional and incredible accomplishment.
Could you replace it with something better designed?
No. No, you cannot. You can in principle design something better, but that's a completely different, quixotic, and useless task.
It's also far from impossible to implement Unicode 'correctly', folks not only can, but do, routinely. It's extensively well documented, there's example code, it's just work.
Also, if your game plan for Unicode-D includes removing the most beloved and consistently demanded feature, emoji: then no, that person in particular is not capable even in principle of designing something better. That game has been lost before it began.
It isn't (and never can be).
> Could you replace it with something better designed? No. No, you cannot. You can in principle design something better,
Something that some people fail to consider, is that one character set is not suitable for all purposes. Unicode is not very good for most purposes though. I think Extended TRON Code has many advantages, although trying to use Extended TRON Code (or some other alternative) for everything would be almost as bad as using Unicode for everything, but in different ways.
> Also, if your game plan for Unicode-D includes removing the most beloved and consistently demanded feature, emoji
I think that colourful emoji should not belong in the character set for text. I also do not want colourful emoji on my computer.