Unicode Character “;” (U+037E)
compart.com
compart.com
This might be case where I'd actually prefer a box, so that I am a little less confuzzled.
My terminal should support only my local limited characterset. Rich rendering should require piping through a specific interpreter whose output is strictly limited to the present terminal (that is, stdout is not redirectable).
I rarely express them within a file. (I do use vim's digraph feature to some extent.)
In contexts where mathematical relations are necessary and used, there are virtually always ASCII-based inputs for them. (I mostly encounter these in GNU units and bc, though a More Serious Numbers person might make use of python, Octave, or Mathematica.)
"A specific rich-text experience" would include contexts such as a Web browser or LaTeX ouput. Note that under most circumstances characters can be input by other and unambiguous methods.
Generally, the Greek alphabet doesn't have homoglyphs that would be confused with Roman. The same cannot be said for other alphabets, notably Cyrillic, say.
The ability to constrain programmatic contexts, or even a web browser's navigaion bar, to specific charactersets, say the local characterset of the user plus 7-bit ASCII (for all international access) seems like it might have strong arguments.
The fact that this cannot faithfully represent all proper names or references in their original language (rather than the user's) is a feature, not a bug.
See "homoglyph attacks" or "homograph attacks":
https://cisomag.eccouncil.org/homoglyph-attacks/
https://blog.malwarebytes.com/101/2017/10/out-of-character-h...
Anyway, according to your view the proper glyph set to display/character set to type is probably related to or based on the set typesetters used in your country 50-100 years ago, during the age of lead type. Right?
(As a side node, you might as well forget homograph attacks. They seemed to be a threat a couple of decades ago, but we've now seen that the bad guys don't bother. They just register some domains and copy logo JPEGs. They send mail from postbanksecurityteam.ru or postbanksupport123@hotmail.com with the Postbank logo in the message, and people fall for it, no homographs needed.)
You seem to be pursuing a similar track here, in that there seems to be a model of me you've constructed and are arguing against.
That model is entirely unfamiliar to me.
If instead you were to ask, say, "what are your concerns and objections to very rich or extensive charctersets, and where do those objections apply", you might find that the discussion is somewhat more productive.
I've hinted at this, rather loudly in my view though you seem to have entirely missed the cue, at the end my earlier comment to which you're replying: homoglyphs and homographs.
Аnd had you done so, I might have answered somewhat like this:
In print, or in writing, a letterform exists as itself. А letter has only one aspect, it's shape. And that shape, in context, defines how that character is interpreted. You don't even have to venture from the АSCII characterset to experience this: I, l, 1, and 7 (in some cases), 0 and O, u and v, uu and vv (or UU and VV), rn and m, fl and A, cl and d, lo and b, Cl and P, S and 5, 7 and T, can be confused, either easily or at least under some circumstances, several especially in handwriting or as ligatures. Previously, ſ and f (the long-ess and eff characters). On some printing devices (notably typewriters), the characters are literally the same, and even some punctuation are constructed of other characters (e.g., bangs, '!', entered as .<backspace>'). You'll find this reflected in certain Roman-alphabet based nomenclatures, with the Library of Congress Classification System being one: it omits the letters I, O, and W on the basis that they're readily confused with other characters (1, 0, and VV, respectively). This is similarly a frequent cause of confusion in serial numbers and other cases where alphanumeric characters are used without regard to standard rules of grammar.
Аnd that's based only on appearance and confusion of characters themselves.
In computers, glyphs and the characters they represent have a dual existence. There is the presentation of the glyph, and there is the machine encoding (in bytes). Again, we don't have to go beyond 7-bit АSCII to encounter this problem. In passwords and serial codes, distinguishing these characters is a notorious source of error, especially when transcribing via hardcopy (handwritten or printed) forms. The issue is that a closed loop might represent a nil or the letter oh, and a vertical stroke any of one, capital eye, or lowercase el. The presentation and the machine coding are at odds.
Аnd that's with just a 128 element characterset. The problematic cases are few enough to enumerate clearly and remember. I do consciously avoid use of the characters mentioned in contexts where they're likely to cause confusion (most especially passwords), and take care to distinguish these characters in my own writing. Not always successfully.
Unicode currently has somewhat north of 144,000 character glyphs represented. That's a larger glyph set than most people's language vocabulary (on the order of 20k -- 40k words or so). And there are a considerable number of similar-appearing glyphs, not only among letters but also punctuation marks (as in the submission here). Not only will I never type these, but I recognise only a very small fraction of them. Аlthough with luck, I might be able to tell when a stray glyph outside my usual working set has snuck into a digital document, there's a strong likelihood that I would be completely oblivious to the fact, as would any other user. Quite possibly yourself as well, Аrnt.
The richness of the Unicode glyph set conflicts with the possible confusion (and malice) which can result.
Many writing systems make use of a small set of glyphs, typically on the order of about 15--50 or so, most especially among phonetic writing systems. Syllabaries and ideographic writing systems (notably Hanzi and Kana) have thousands to tens of thousands (typically 2,000--3,000 are in wide use).
Аnd in contexts where confusion can be somewhere between confusing, annoying, or outright harmful, the flexibility provided collides with potential for harm.
There are places in which richness is deserved and warranted. Аnd there are places in which it is not. What I've proposed is that most tools that users use in directly interacting with their computers --- command shells, navigation bars, input dialogues, code editors, and the like --- are best limited to only the glyph sets which are most immediately useful to that user. That is, 1) their own native characterset), 2) that which has become the de-facto standard for computer encoding (127 bit АSCII), and possibly 3) a small set of additional useful characters (of which the Greek alphabet would be a reasonably safe bet, though I'd argue that neither its punctuation glyphs (the subject of this submission), nor other easily-confused charactersets (e.g., Cyrillic) would not be unless those fell under the first condition above.
Elsewhere, in clearly formatted and typeset contexts, with an appropriate language specification, other charactersets migh be used. Keep in mind that as a service to the user or reader* those charactersets offer minimal value (they're quite literally foreign), and significant potential for harm.
(In case you'd missed it, every instance of an uppercase 'A' in this comment has actually been a Cyrillic 'А'. A case in which Аristotle's law of identity is literally not true: А is not A.)
It's fair game to ask me to memorize ASCII. It's the goddamn alphabet, we can safely assume first grade level education, but I don't want characters that I don't recognize at a glance and I can't easily type in my code/output/whatever.