The revenge of Unicode
eclecticlight.co
eclecticlight.co
And of course there certainly are Hungarians that also write in Thai. So what about mixed Hungarian/Thai text? Should that be a multiple (interleaved) encoding text? I mean as an alternative to “149,813 ‘characters’” that “any one human is [not] likely to use [most of]”.
Unicode is a great thing and has made computing more accessible to those who read and write in languages that don’t use the Latin alphabet.
Except, of course, anyone relying on a screen reader or other accessibility tool.
https://kence.org/2020/10/02/accessibility-of-unicode-as-sty...
[1] https://util.unicode.org/UnicodeJsps/list-unicodeset.jsp?a=[...:]
https://github.com/wanderingstan/Confusables
From the readme:
E.g. "℮1೦" would match "Hello"
"Hello" gets turned into the following regex of character classes:
[HHℋℌℍΗⲎНᎻᕼꓧ𐋏ⱧҢĦӉӇ]
[e℮eℯⅇꬲеҽɇҿ]
…
Edit: interesting that some versions of the letters are already being filtered by the commenting system here.Chinese sends it's regards.
But also text interchange is useful even if most people only use a fraction of it most of the time. How else is an English document supposed to quote Chinese characters for instance (this happens all the time on wikipedia pages which share the original name of something next to the translation).
The days before unicode were dark days, nowadays besides the ocassional oddity text pretty much just works without users thinking about it -- thanks to unicode.
In other news. How many Hungarians are ever going to be writing in Thai?
EDIT: Used the wrong quote here. Fixed.
> All those ‘characters’ enable deliberate misuse, where visual similarities are exploited to spoof people over identity or worse.
Communities like mathematicians are more to blame for that than Unicode.
Unless maybe you want a Latin script/Cyrilic/Greek etc. unification (for the relevant characters).
> We still do a great deal in life using text that can be searched rapidly and readily. Sometimes it pays to obfuscate that so that only humans reading it will understand what it says.
I guess that would work in the same way that rot13 encoding (like spoilers) does; not worth to decode since only a few specially interested people do it. But it’s not like it’s difficult if anyone puts some effort in. But yeah, such nerdisms can definitely be useful as long as they remain a subcultural thing (presumably always will be).
The point is that Unicode has to deal with how text has been used by various communities in order to be the text encoding standard to rule them all.
[1] Unlike in prose where bold and such is just for emphasis. And at worst you can use informal markup (similar to MarkDown) and convey the same thing.
> Like the symbol for infinity is from the Hebrew alphabet.
Are you perhaps not confusing infinity with the symbol used to denote the cardinality of infinite sets? If the infinity symbol really comes from the Hebrew alphabet I'd love to hear more.