> Who (process-wise) is responsible for converting bytes to pixels?
The operating system, with minor exceptions (word processors, for example). Rendering logic is too complicated to be embedded into every application. And you get better accessibility and consistency.
> How do users on social media put in their name? How is it stored?
Keyboards, or their preferred input methods, and stored in UTF-8. I'm not saying to get rid of all strings, just don't use it for infrastructure.
> How do users get urls with specific usernames?
You don't, because that's how you get little Bobby FRACTION-SLASH. Also, if you have a valuable namespace like URLs, people will hack each other to get valuable names, but that's only tangentially related.
> Now take all that and multiply by the complexity of world languages, many which don't even map to one glyph == one morpheme. The ol' apple message crash bug was due to the property of some Arabic not being monotonic in rendering space vs string length.
That's exactly my point! You get this multiplied complexity when people try to peek into string contents instead of treating them like black boxes. Stop with the dangerous string operations and you now support usernames with zalgo-ed hieroglyphs if that's what users want.
> I think we could have skipped utf8 and just gone to 4byte runes. But even then, that would not have avoided the above bug.
> Utf16 is a hot mess though, worst of all worlds.
Agreed with UTF-16. I like UTF-8, and I honestly think it solved our encoding problems for non-legacy applications. Everyone should be using it, as long as the contents are for human consumption only.