I suppose this is inevitable when you tasked with representing literally every symbol in existence. You couldn't pay me enough to touch this problem with a ten foot pole (this and text rendering).
I suppose this is inevitable when you tasked with representing literally every symbol in existence. You couldn't pay me enough to touch this problem with a ten foot pole (this and text rendering).
Writing seems simple — children do it routinely — but like a biological system it evolved over millennia in a ton of different directions. It’s coupled with emotional, practical, and even, yes, moral issues that operate on both deeply personal and social issues. This is hard to capture in software.
Unicode made a couple of hard decisions right up front. I hate them but they were smart and Unicode would not have survived had they not made them. One was round trip with legacy character sets, which meant encoding a lot of redundant characters (English and German “A” have he same code point, but Greek “A” and Russian “A” do not, nor does an “A” that appears in a Japanese code table. Second was abandoning attempts at Han unification, which had its own linguistic, emotional and political issues.
People are complicated and so are their languages so wrestling the whole thing into a tractable system has been worth the effort.
Huh? Han unification happened.
Imagine if Unicode has to start dealing with the temporal change that for example the Olson TZ database [2] has to!
[1] https://shkspr.mobi/blog/2019/06/quirks-and-limitations-of-e...
One example of how this is a huge clusterfuck is that until recently, Windows Notepad opened and saved everything with the Win-1252 encoding scheme (labeled as ASCII in the app). The web, and the other popular OSes, on the other hand, are standardized around UTF-8. So if you download a txt file from the web or OS without a BOM, and you open it in Notepad, you can get characters that looked right in your browser, but not in Notepad.
There are smart algorithms out there that can detect character encoding pretty well, but none of them are perfect (as far as I know).
The Win-1252 default and the fact that most computer users have no idea about character encoding have caused all sorts of headaches for me with the reporting software I work on.
The idea that computers should support cultures, and not the other way around, is pretty recent.
Gruß, stkdump
Um, no. The words were originally written that way. Ä, ö, ü and ß actually developed from ligatures for ae, oe, ue and ss, long before computers were a thing.
I modified your first sentence to make it more generic and applicable to many other things in software.
Barring that, we could have all used UTF-8, but windows really screwed that up, and none of the arguments for 16-bit alignment vs 8-bit alignment for processing really hold water.