Falsehoods Programmers Believe About Plain Text (2021)
jeremyhussell.blogspot.com
jeremyhussell.blogspot.com
“Text copied from MS word and pasted on a text field will have some resemblance to the original text”
I don't know how you can fuck up like this.
Characters are integers, that is a fact. Most of the time, they're a byte, and to be honest, English/ASCII is the most important text setup for computers and it's fine to only support that. For the niche foreign edge cases, Windows supports UCS-2.
So for example if you are writing a CLI utility for the European market, you could probably get away with doing very little, it would be an ugly experience for, say, German users, but it would be quick. You could go to the next level, eg, supporting umlauts and ß, and that would have a big payoff. You could probably stop there. But if you were writing a utility to do OCR on ancient texts, you'd have to go full-bore and support everything..and you'd still be wrong on occasion.
It's just a question of economic tradeoffs being explicitly judged and thought about, and not assumed away.