Unicode in URLs and usernames and filenames is just so easy to trick people with.
Will there be Unicode email addresses?
Unicode in URLs and usernames and filenames is just so easy to trick people with.
Will there be Unicode email addresses?
Apparently, yes. [0]
> making things more equitable for other languages and cultures
Well, yes. People often hate how compilcated Unicode is, but they also tend to forget that even ASCII was not sufficient for writing English in a proper way. I am not by any means old, but even I still remember the period on the web when non-Unicode encodings were relatively common and how problematic it was.
And, with assumption that by “other cultures and languages” you mean non-English speaking regions: Tolerate is a wrong word. Most people are not English native speakers. If we would go the route of deciding who is tolerating who, nations using Chinese Hanzi and its derivatives would be probably the first to claim they tolerate all the others by the speakers’ headcount alone.
Moreover, deciding what is/can be executed is not so obvious in my opinion. PDFs can have JavaScript embedded, SVG too. Flash games? Python scripts on computers with python installed? It is not just the names, but the very knowledge of what can be considered “executable.”
I know of both the over-strike method for traditional typewriters and extended ASCII. But I have no knowlede of a ASCII (extended or not) supporting every sign commonly used in English. Granted, it is not as significant as other languages missing whole letters, but signs like ”,“,—,–,’, which are often replaced with simpler ASCII alternative (",') or their combination (---,--).
Over-strike, from what I remember wasn’t implemented in any well-spread way with standalone ASCII, although I could be simply not aware of it. It gives us then a few more characters rarely used in English language in loanwords (née, naïve, façade).
Skipping diacritics with uppercase letters on displays with too little space happens with Polish language too, so I am aware of the practice.
In any case, yes, overstrike is just an approximation, but believe it or not, ASCII was designed to make it possible. And yes, it kinda sucks, and yes, it doesn't really make ASCII a multi-byte encoding, not exactly. But it's a funny thought that ASCII kinda almost was a multi-byte encoding. One can imagine ASCII evolving to treat BS<mark> as not unlike Unicode combining codepoints to make it possible to get proper diacritics even on capital letters, but ASCII would still have been a dead end, naturally.
The comment you're replying to just said ASCII without qualification. By definition this is not multibyte, or even whole byte (assuming by byte you mean octet). ASCII is a 7 bit standard, anything else is some other encoding.
I'm quite certain that ASCII was designed to make overstrike feasible. Overstrike just wasn't explicitly part of the standard. The quip about it being a "multi-byte encoding for Latin" was a joke made funnier (to me, and perhaps only to me) by having a kernel of truth in it.
This guy made a whole business out of it:
Let me rewrite that from the opposite perspective.
The reason I use it is so that I can write in my language.
Writing in a language the reader doesn't understand is ... not so fine.
Writing in a way to give the appearance of one message but the machine-recognised existence of another is ... wrong, malicious, and harmful.
At root the issue is that encodings and graphical presentations aren't the same thing. 7-bit ASCII is limited and constrained, but as a universally understood encoding those specific characteristics are useful benefits. Yes, it means that representations are limited. But that's the essential trade-off for a lack of ambiguity.
And even within ASCII, there are homoglyphs or near-homoglyphs: {0O,1lI, 5S} being the most frequently encountered. Kerning and ligatures may present others, as with {m, rn}. In historical documents, distinguishing {ſ, f}.
And, yes, of course, email addresses in local languages, why the heck not? People in the world want to use their language! And yeah, even Hebrew, Tibetan, or Mongolian (in vertical script).
That comment of yours reads to me like an lazy post of a unilingual or uniscriptal cultural imperialist who really does not care about other languages.
Your attitude makes me rather angry, because Unicode is such an amazing achievement for the world, and your comment just comes off totally ignorant of that. It should be clear that combining the world's languages into a common standard is really, really difficult and will inevitably create something that is much more complex than your beloved ASCII. And it should also be clear that in such a complex standard, you cannot (a) solve all problems the first time you try, (b) you cannot try a second time, because no-one will adopt yet another such standard.
So just read those Unicode documents in order to understand. Those people really try and there are security consideration documents, and they are extended all the time. And take those new security warnings for what they are: they are problems with a complex system. It's expected. So when they get known, react calmly and figure out whether you need to fix anything.
well ok, it seems to me that it is more likely that in most cases Unicode probably has no effect on security one way or another, but there are a few edge cases where it does. Here is one case where using a particular character in a filename led to problems 8 years ago in an operating system historically known for not having the greatest security.
We can see here also a clear example of the edge case theory, almost every character in every language supported by Windows would not have caused any sort of problem when used in a filename. But there are a few characters where they would, this is probably an example of things programmers think they know about unicode or language or whatever, because we go around thinking that each unicode code point just represents a character in a language but some encodings have ways to represent little weird behaviors of particular languages, for example interlinear annotation characters https://www.unicode.org/charts/nameslist/n_FFF0.html and it is generally weird behaviors that end up being security hazards (not sure if there is any security hazard in interlinear annotation characters but wouldn't be surprised)
Are you overreacting? Unless you were proposing we rip it out, no, you're not.
Unicode isn't going anywhere. That's because the scripts that human languages use aren't going anywhere. And people do mix scripts, too.
There are answers though. First, Unicode has been a learning experience, and we're still all learning. Second, one of the outputs of all that learning that has happened so far is UTR #36 (https://www.unicode.org/reports/tr36/), which does cover a lot of these things.
I mean, obXkcd: https://xkcd.com/327/
But spelling out what you had in mind would be helpful here.
Still, a limited set.
Yes, errors are made and occur. They're reasonably easy to code defensively against.
Unicode ... vastly expands the attack interface.
> Unicode ... vastly expands the attack interface.
Limit yourself to 0..9 to be safe.
And how that compares to the number of 7-bit ASCII characters?
And how many special cases would have to be considered?
Scale matters.
The risks:reward ratio from 7-bit ASCII is low and manageable. The expressive capability is high. No, it's not a perfect representation for all languages. It is, however, a sufficient one, where common understanding is necessary.
At 128 vs. 10^20 possible codepoints, many in Unicode with side effects, the problem of deceptive or unintended use is exceedingly high in Unicode.
With far too many special cases to hold in human working memory, or even reasonably within most code-bases.
Well, that's a lot like saying "software, in general, is just a giant security flaw we tolerate because it does useful things". It's kinda the point aint it?
In 21st century, even Americans aren't able to write their own names and place names with pure ASCII anymore. So Unicode is there from necessity, not as an addon.
All technologies have intended / desireable, and unintended / undesireable effects.
The more powerful and flexible a technology, the more likely there are unintended / undesireable effects.
Many of those unintended / undesireable effects are themselves not obviously apparent, not immediately manifest, or both. All of which makes risk assessment all the more difficult.
Software, in that sense, is inherently a security flaw, as it virtually always manifests unintended and/or undesireable effects.
This isn't necessarily an argument to ban all software, though the Butlerian Jihad are taking notes. It is an argument, however, for acknowledging risks, raising awareness of them, and taking reasonable steps to guard against them.
C's flaws lead to the development of several different competing programming languages. Various flaws in cryptography lead to the development of algorithms that achieve the same result, but have far fewer implementation gotchas. IMO, the time has come for the development of a truly strict variant of Unicode that still supports the primary objective, but learns from Unicode's mistakes.
The problem is that nobody cared. Browsers invented punycode instead of following tr39, email ditto. But ok, at least something. Java did it, cperl did, rust did it.
Everybody else is vulnerable. Esp. most other programming languages, filesystems and login systems. https://github.com/rurban/libu8ident/blob/master/doc/c11.md
So you don't know about mailoji.com? It was on HN a couple months ago I believe.
Same thing really.
However, tangential to the grandparent's topic, I do imagine there are very specialized cases where Unicode is pretty objectively unnecessary.
There are certainly contexts in which Unicode unambiguously and demonstrably leads to security weaknesses and issues. See generally homoglyph attacks.
At the heart of the lie and damage is the existence of a message which appears to say one thing but in fact says something different. It's the very limited nature of 7-bit ASCII, 128 characters in total, which provide its utility here. Yes, this means that texts in other languages must be represented by transliterations and approximations. That's ... simply a necessary trade-off.
We see this in other domains, in which for the purposes of reducing ambiguity and emphasizing clarity standardisation is adopted.
Internationally, air traffic control communications occur in English, and aircraft navigation uses feet (altitude) and nautical miles (dstance) units.
Through the early 20th century, the language of diplomacy was French. The language of much scientific discourse, particularly in physics, was German. And for the Catholic Church, Latin was abandoned for mass only in the 1960s.
Trading and maritime cultures tend to creat pidgin languages --- common amongst participants, but foreign to all, as distinguished from a creole, an amalgam language with native speakers.
A key problem with computers is that the encodings used to create visual glyphs and the glyphs themselves are two distinct entities, and there can be a tremendous amount of ambiguity and confusion over similarly-appearing characters. Or, in many cases, glyphs cannot be represented at all.
Where the full expressive value of language is required --- within texts, in descriptive fields, and in local or native contexts, I'm ... mostly ... open to Unicode (though it can still present problems).
Where what is foremost in functionality is broad and universal understanding, selectinga small standardised and widely-recognised characterset has tremendous value, and no amount of emotive shaming changes that fact.
As an example, OpenStreetMap generally represents local place names in local language and charactersets. This may preserve respect or integrity to the local culture. As a user of the map, however, not knowing that language or charcterset, it is utterly useless to me. Or, quite frankly, anyone not specifically literate in that language and writing system.
It's worth considering that the characterset and language in question are themselves, adoptions and impositions: English was brought into Britain by invaders, the alphabet used itself is Roman, based on Greek and originally Phoenecian glyphs. English has adopted or incorporated terms from a huge set of other languages (rendering its own internal consistency ... low ... and making it confusing to learn).
International communications and signage, at airports, on roadways, in public buildings, on electronic devices, aims at small message sets and consistent, widely-recognised symbols, shapes, fonts, and colours. That is a context in which the freedoms of unfettered Unicode adoption are in fact hazardous.
(Yes, many of those symbols now have Unicode code points. It is the symbol set and glyph set which is constrained in public usage.)
And the simple fact is that a widely recognised encoding system will most often reflect on some power structure or hierarchy, as that's how these encodings become known --- English, Roman Alphabet, French, German, Latin, etc. Small minor powers tend not to find their writing systems widely adopted (yes, there are exceptions: use of Greek within the Roman empire, Hindu numbering systems). Again, exceptions.