I appreciate your reply, which I never expected in this situation.
If my understanding is correct, your thesis is that Unicode should be hidden from application programmers as much as possible, much like the fact that GC hides memory management so to say. Not to say Unicode is bad or even shouldn't exist at all, but something like that it has to be abstracted away. This is a much more reasonable than most (quote-unquote) Unicode criticisms indeed. I'm not sure whether this is possible in the near future however, for reasons I'll work out here.
----
From perspectives of API consumers, most if not all programming languages have a suboptimal design for the human text. In fact the type name "string" itself is inappropriate, its name comes from the assumption that a human text is a string of symbols, which is not incorrect but not helpful either. A proper "human text" type (or a collection of them) should ideally be able to do the following:
- An abillity to contain additional linguistic informations like locales, grammatical genders or numbers if possible. Some can be guessed, some can be retrieved from external contexts (e.g. HTTP `Accept-Language`), some have to be retrived with a consent. Any text operation should retain them if the corresponding text is also retained.
- A language-aware formatter. For example `"Total: ${n} files"` should automatically change "files" to "file" when `n` is 1. Moreover, `"Total: ${n} ${objectName}"` should do the same if `objectName` is an English text "files". (This is why every text should retain linguistic informations!) Of course the format text should be translatable (say, to "파일 총 ${n}개" in Korean) and that shouldn't change the original code.
- Proper textual isolation. If my text is composed of multiple scripts or languages, they should not affect each other in any way, and should be displayed in the best way possible. For example a missing font should not give broken boxes; either the font should be downloaded on demand, or a note about missing font should be shown instead. Inserting RTL texts into LTR texts should not flip either of texts (unless it is required by the surrounding languages). Basically, no surprises even if you don't know about them.
- Situation-aware alternatives. Even after the formatting, a long text that doesn't fit into the UI should be shortened in the way that as many information is preserved as possible. For example the text "Nice to meet you, ImagineAVeryLongUserNameHere!" will be cut into "Nice to meet you, ImagineAVery..." today, but one should be able to turn this into "Hello, ImagineAVeryLongUserNam..." from the formatting layer.
None of these operations actually concern Unicode, but they are incredibly hard---if not impossible---to build. Most of them are at best fuzzily defined or often undefinable. So we are left with a number of localization and internationalization libraries which are ignored by most developers to say the least. A mere "string" type is a norm, a program has no idea about the text and proceeds with faulty assumptions, and users are so accustomed to bad text handling that they even don't expect much. If enough users complain there is a chance of improvements, but even that is done by a case-by-case basis.
---
Unicode sits at the level much lower than what I've imagined before. It is not even a component to build the human text. It is a component to build a string, that can be somehow used to build the human text if one is very careful. Most human text operations can't be done with Unicode alone.
For example, people argue that the number of "characters" in a single Emoji sequence (say, one mentioned in https://hsivonen.fi/string-length/) should be 1 and others don't make sense. This is meaningless because it will appear as a number of broken boxes if emoji fonts are not installed anyway (but not five, because it contains two default-ignorable code points). It matters what the number of "characters" is used for, and that's a whole point of the linked article. And the definition of user-perceived characters does vary over locales, so you can't count them without a linguistic information anyway.
You may still argue that Unicode algorithms are designed for the human text encoded in strings. That's a very nuanced argument, because one can also argue that they are the best effort approximation of human text operations for strings, in which case they are not the human text operations themselves. For example many languages have a case conversion operation over strings, with a varying degree of Unicode conformance (ASCII-only, simple fold, full fold, locale-dependent fold, title case, ...). But the case conversion itself is not the human text operation! Even assuming bicameral scripts, some texts are never capitalized (e.g. "McDonald" frequently capitalizes to "McDONALD", not "MCDONALD"). The human text operation, here full capitalization, needs much more than the Unicode case conversion algorithm.
Given this, it is a misguided effort to make strings more aligned with the human text, because it is not possible at all. As people frequently mistake strings as human texts however, the second best thing is to get rid of any string operation. Swift almost did this but retained a default grapheme view---I think it is actually worse given the instability of (extended) grapheme clusters over time, but also understand why they had to do that. [1] The third best thing is probably to stress that a string is not a human text, in the same way that a floating point number is not a real number. And the original post, in spite of some errors, did a good enough job in this regard.
[1] There is also a precedent of Raku's NFG (which dynamically allocates a negative code point for new grapheme clusters seen), but this is more or less an optimization of the graphemes view. The current Unicode has an infinite number of distinct grapheme clusters by design.