At some point we have to at least gently blame the overall developer community. We've had actual decades to figure this out. Joel's
The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets (No Excuses!)[1] manifesto was written almost 20 years ago!
99% of programmers shouldn't be treating text as anything but opaque blobs of data that, when sent to a rendering function, become readable by users. Don't try to take text blobs apart and make assumptions about their binary structure. We should not be looking under the hood and trying to divine things like how many letters are in them (what's a letter?) or trying to manually transform them in some way (case, splitting, concatenation, input sanitization and so on). Use tried and true libraries for these. Same for when you need to transform it into some particular encoding. The only safe question to ask of text is "how much memory does this thing use?"
It is truly a minefield, but it's a minefield that developers are expected to understand and navigate. "Text in different encodings are scary outliers that we shouldn't support" is pee-wee league level software engineering.
1: https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...