IDN is crazy
daniel.haxx.se
daniel.haxx.se
I feel IDNs are part of a larger trend of "Unicode should be everywhere", without any consideration for whether this is actually a good idea. Programming languages are guilty of this too, with many allowing almost any Unicode codepoint in identifiers. This has created an entirely new class of security vulnerabilities that are serious enough that VSCode now flags such codepoints by default if they look similar to certain ASCII characters.
Basic ASCII Latin characters are the only characters that can be entered using almost any input device in existence without additional tools or configuration. This makes them a universal baseline, and software should carefully consider whether deviating from that baseline is really worth the security vulnerabilities and incompatibilities it may introduce, just so people can write their source code comments in Vedic Sanskrit, for a programming language where every keyword and almost every function name is in English.
Heck, it might even be a reasonable idea to bake anti-unicode into the system (kernel) and require elevated access to process it at all in a sandbox...
But that doesn't mean domain names should, or identifiers in programming languages. Those are critical, fairly low-level technologies, and allowing all of Unicode raises significant security concerns. I strongly believe those concerns outweigh the usefulness of non-Latin characters in those contexts.
The accusation of anglocentrism is unwarranted here. Every mainstream programming language ever created has all of its keywords based on the English language, including languages like Ruby and Python whose creators aren't native English speakers. This isn't anglocentric, it's a best-effort attempt to make the language as universal as possible in the world we live in. Internationalization at this level of technology simply leads to fragmentation. Domain names and programming languages are very different from user-focused free-form web content.
Also, I disagree that domain names aren't user focused. It's the first thing you learn as a regular web user (at least that used to be the case before, now it's usually the app name).
Well, now you have. I'm not a native English speaker and I'm not from a country where English has any official status.
is that a joke? or are you really under the impression that most languages use subsets of the Latin alphabet? even English uses letters that aren't in ASCII, much less literally every other one.
Also, “Those two dots, often mistaken for an umlaut, are actually a diaeresis… Most of the English-speaking world finds the diaeresis inessential. Even Fowler, of Fowler’s “Modern English Usage,” says that the diaeresis “is in English an obsolescent symbol.””
https://www.newyorker.com/culture/culture-desk/the-curse-of-...
The main argument against is that everybody is accustomed to the current situation and knows that when a brand contains diacritic marks, they should be removed when entering URL. If IDN is introduced, it would be confusing as some would use ASCII version, some IDN version, while most would just buy two domains instead of one.
But it doesn't, so that's an irrelevant hypothetical.
Which is why almost every institution everywhere uses kilometers or miles to measure distances, not parasangs or barids.
Well, the way the world works created IDN and Unicode identifiers in languages. How about that?
Wait, IDNs exist, they are "the way the world works". You are arguing that they should not exist, because you don't have a use for them and they inconvenience you. How can you justify your desire for the world to work differently by saying that it's "the way the world works"?
Are palindromes they key?
Examples of Issues Found with RFC 3454
4.1. Dhivehi
Dhivehi, the official language of the Maldives, is written with the
Thaana script. This script displays some of the characteristics of
the Arabic script, including its directional properties, and the
indication of vowels by the diacritical marking of consonantal base
characters. This marking is obligatory, and both two consecutive
vowels and syllable-final consonants are indicated with unvoiced
combining marks. Every Dhivehi word therefore ends with a combining
mark.Perhaps with current level and awareness of multi-lingualism in computer science/engineering, this whole discussions about anglocentrism in computing and possible solutions is too early to take place; very few are even aware, much less accept, that the spoken language active during a decision making affects its result for multi-lingual people.
What do you gain from IDN domain names other that nobody who doesn't speak your language can't even type or remember them? Apart from the obvious security issues mentioned in TFA. Same goes for handles or ids. Don't allow non-ascii for usernames to avoid scams. If you insist, have a separate display name that's displayed alongside.
That people who do speak the language can type them?
Every keyboard layout in existence for languages not based on the latin alphabet has a trivial way to switch to Latin input.
It's a reminder that we live in a shared world.
And if we start seeing more keywords in alternative scripts...
That's more improvement in strong typing, IDEs, LSP, dev containers, static analysis and a slew of related technologies. It's exhausting, but it seems a little more fair than the URL debate. Of course there will be human auditing and approval advancements there that are exploitable. But it is what it is.
Also with a push from Google (and Firefox too) most users no longer type URLs - they enter a search query and go to a site from search result. I don't like this trend but site names become increasingly invisible for users.
Also, Firefox sadly has all individual countries TLDs enabled by default if I enable the whitelist. It should be the opposite.
As for the homograph problem mentioned in the article, that's not an anglo-centric fluff. It's an actual major problem that affects speakers of all languages. It hinders the ability of people to reliably type or recognize domain names. That's going to cost some people their life savings.
Overall, IDNs would cause more harm than good by making URLs unreliable. It would just serve to empower search engines.
The first part of this sentence does not imply the second. The world's first programmable computer was in fact created by a German. That doesn't make German the language of computing.
> are you okay with only using Cyrillic as the sole script
No, because Russian/Cyrillic lacks the properties that make English/Latin so useful as a universal language/script:
* Almost all languages have widely used, standardized transliteration systems into a Latin script. That's not true for Cyrillic.
* English is far more widely spoken and read than any language written in Cyrillic.
* The subset of the Latin script used by English is essentially the common denominator of all languages written in Latin. That's an extremely important property and immediately predisposes English for the role it has today.
There are objective arguments for English as a universal language for computing. This isn't just "cultural imperialism" or something. English is not my first language, but if I try to imagine my first language being used in its place, I immediately see dozens of serious issues that would arise, all of which are non-issues with English.
> * Almost all languages have widely used, standardized transliteration systems into a Latin script. That's not true for Cyrillic.
There may be standard transliteration systems for non-English alphabet languages, but there are often multiple standard transliteration systems (e.g. Russian, Arabic). Also, Chinese transliteration is pretty gobbledygook without diacritics that aren't present in the English alphabet (and kind of gobbledygook even with diacritics).
> * English is far more widely spoken and read than any language written in Cyrillic.
What if the de facto language of computing was Mandarin instead?
> * The subset of the Latin script used by English is essentially the common denominator of all languages written in Latin. That's an extremely important property and immediately predisposes English for the role it has today.
The Latin script used by English is exactly all the letters you need to write English. Many other languages written in Latin script either have more (e.g. Spanish, Norwegian), or fewer (e.g. Italian, Serbian), or have different rules for certain things (e.g. Turkish).
Mandarin isn't "widely" spoken. It's spoken by a large number of people, almost all of which reside in a single country. This makes it utterly useless as an international standard.
And the Chinese writing system would have been far too complicated to become the foundation of computing. We're talking a couple dozen characters vs. tens of thousands. Only recently did it even become possible to accurately represent the whole gamut of Chinese writing using digital systems.
That said, technology follows the needs of people, not the other way round. If we were living in a parallel reality where Mandarin had the cultural cachet of English, I'm positive we would have immediately subjugated computers to deal with Chinese characters ;-)
The norm for international communication is that zero people on either side speak the language. Internal communication anywhere in Achaemenid Persia took place in Imperial Aramaic, which was spoken by some subject peoples in the west of the empire.
International communication between Mitanni (speaking an unnamed Indic language) and Egypt (speaking (Afroasiatic) Middle or Late Egyptian) took place in (Semitic) Akkadian.
International communication between Japan and Korea took place in classical Chinese, spoken by neither side and unrelated to the languages spoken in either. Similarly, international communication between Vietnam and China took place in classical Chinese, spoken by neither side (though closely related to the languages spoken in China).
Treaties between early modern Russia and China could not be concluded in classical Chinese, so they were concluded in Latin, spoken by neither side.
It just isn't an expectation that international communication will use a language that is spoken by any party to the communication. This is for the obvious reason that anyone who wants to join the conversation uses whatever is already in use, even if the language currently in use has been dead for a thousand years.
Only true in the sense that that is how English spelling is currently defined. If you wanted to come up with a script for English from scratch, it would look nothing like the Latin script. It would not be isomorphic to the Latin script.
Compare this inventory of phonemes in General American English, with 24 consonants and 13-15 vowels, of which only 5 are diphthongs: https://en.wikipedia.org/wiki/General_American_English#Phono...
Well, it's not true by any measure. It is not the case that almost all languages have widely used, standardized transliteration systems into a Latinate script.
> The subset of the Latin script used by English is essentially the common denominator of all languages written in Latin.
This isn't true either. Why would an Italian consider the weird English letter 'j' more universal than the weird French letter 'ç'? The French might not mind 'j', but what are they supposed to think about 'w'?
The "objective arguments" you're saying is simply hand-waving on the fact that the English (and Americans, which mainly descended from English pioneers) have spoken English. I should remind you that French is still the primary language of diplomacy (even it looks like English has replaced it) and French has also a robust transliteration system to many languages (and in some cases which is even better than how English handles it due to its fewer phonological edge-cases), mainly because it also have historically controlled some parts of Africa, Asia, and the Pacific, and have a serious diplomatic force for areas which the French haven't bothered to conquer with.
Now, the question is simply what if Cyrillic is the lingua franca of the computing world, and you've never answered it directly. You instead doubled-down on the supposed benefits of the standard English and alphabet which is not what I'm asking. You might not have an answer here, being that it's hard to imagine that alternate reality, but I should remind you that a lot of people globally have significantly different scripts and doesn't know any Latin letters. Sure, the Chinese knows it, but it doesn't apply to many other people in Asia and Africa.
Which is in particular?
>Almost all languages have widely used, standardized transliteration systems into a Latin script. That's not true for Cyrillic.
Combination of historical circumstances and you confusing cause and effect. You can write english in cyrilic.
>English is far more widely spoken and read than any language written in Cyrillic.
See above.
>The subset of the Latin script used by English is essentially the common denominator of all languages written in Latin. That's an extremely important property and immediately predisposes English for the role it has today.
Why? It makes english almost the only language you can write properly. This is a bad property.
The attack is someone you don't trust modifies your source code, using unicode to make the modifications more subtle. What type of threat model is that? There's tons of ways for a malicious party to sabotage source code. Most are probably more subtle than unicode hacks. It is like complaining about how the person you gave a key to your house to might rob you.
I do think that in a code editor already doing syntax highlighting, individual tokens should be bidi isolated so you can't have RLO and friends affecting too much, but beyond that i really think this whole issue is way overhyped.
If the malicious party has unfettered write access to your source code, then sure. But if all changes get reviewed by a maintainer before being merged, then there's way fewer ways, and Unicode attacks are a really large portion of them.
That's not at all how the open-source world works. Many unknown (to the project maintainer) people can create pull requests and try to sneak in a backdoor through, say, an homograph attack.
if ( id == "root" && some_other_condition )
Pull requests comes in: if ( id == "rοοt" && some_other_condition && some_even_better_condition )
I totally vet that! Beautiful added security measure, that "some_even_better_condition" is the nuts, nothing can go wrong: it's an added protection! I merge that ASAP and release!Except your diff tool missed that on that same line "root" was changed to "rοοt" (and the attacker created an account named "rοοt" on your website).
Game over.
The point is not to bitch about my pseudo-code or how account creation do work: the point is that it's easy to miss an homograph attack during a review.
It would sidestep the entire issue if the code properly checked ACL's, role, or userId.
I feel like many diff tools highlight character changes besides line changes. This is especially helpful when someone makes changes to a long line.
As far as input devices have you used an IME hands on before? Non-latin characters aren't really any more "additional tools or configuration" than when during install you select "English" as your language. Others may select something like "中文(中华人民共和国)" and now without additional tools or configuration they are entering via pinyin instead of qwerty even though it's the same physical keyboard symbols being pressed. These IMEs don't always have a 1:1 latin character to unicode character mapping you can rely on either e.g. for the prior example on modern Windows the same set of latin characters can result in different outputs in a similar way to how english autocompletion suggestions on a mobile device work.
That doesn't work; without additional configuration you'll just end up entering the same ascii that matches the keys you press.
You have to enable the IME. (In the example you appear to be suggesting, you do that by hitting shift.) You can't be in Chinese input mode all the time, because that would make it impossible to type non-Chinese text.
Regardless I don't consider this additional configuration anyways in the same way I don't consider holding shift to get capitals changing your input configuration but maybe that's just me.
That's a command; hitting shift by itself to determine your input mode would be better analogized to capslock.
Sure, but with ASCII there are just two or three sets of homographs ("l1I|", "O0" mostly), whereas with Unicode there are potentially thousands of confusable characters, and more are constantly being added.
> As far as input devices have you used an IME hands on before? Non-latin characters aren't really any more "additional tools or configuration" than when during install you select "English" as your language.
This ignores the millions and millions of devices still in use (and not going anywhere) that don't have "input methods" or anything like it.
Try setting up an IME on a computer running DOS (which is still everywhere in many public offices), on an industrial machine with a basic keyboard, on an aircraft avionics system, etc.
Basic English Latin is the common denominator. That won't change until all those "long tail" devices are gone.
Regarding the millions of pre-unicode devices the exact same can be said about public machines running DOS, industrial machines with basic keyboards, aircraft avionics systems, etc of non-English countries too. After all the rest of the world didn't just use latin text or avoid computers until Unicode finally appeared and was supported. Many did, many didn't - the systems and encodings being so painful is what resulted in Unicode.
The whole point of using Unicode for base things like programming languages, domains, etc is so everyone from all languages (including English) can take advantage of the 10s of billions of devices that do support unicode instead of being limited to the long tail of a millions of devices until every last one is turned off. The support of these use legacy use cases was designed into Unicode and is why ASCII maps to the first 8 bytes, Unicode is just smart enough to not require everyone do that for every device because some 30 year old machine in a warehouse needs it.
Add stress, dyslexia, poor eye sight etc on top of those and you can see how this becomes a real issue. Humans simply aren’t very good at string equality.
As for mitigations, in ascii only you could color code lowercase/uppercase/numbers/symbols differently I guess, enforce font that renders glyphs different enough (probably monospace), highlight the domain name, etc. But even these solutions only go so far, and doesn’t naturally translate to all of Unicode or all situations where identifiers are needed.
[1] https://no-hanja-domain.github.io/ (in Korean)
In the end the anglosphere should never be able to dictate the usable charset on the entire internet. That'd be utter nonsense.
I find that suggestion outrageous.
To the topic at hand, why is that outrageous/flame bait? They seem to have pretty sane justifications, whereas most of the rebuttals seem to be focused on presumption of personal details. Not that I expect you to be immediately swayed or anything, but I don’t understand why you wouldn’t categorize that as disagreement rather than flame bait (am I missing some context?)
p-e-w's comment is uncompromising and dismissive. It doesn't seek to find solutions, but ridicules the very idea of someone wanting to use a language other than English in source code and domains.
When challenged, they have pushed the conversation towards source code, a tangent they introduced, rather than IDNs which most of the discussion is focussed on.
The HN guidelines say:
> Eschew flamebait. Avoid generic tangents.
> Please don't pick the most provocative thing in an article or post to complain about in the thread.
The tone of the comment could hardly be less provocative.
(A supermarket receipt on my desk just caught my eye. It says "se åbningstider på www.føtex.dk" at the top.)
FWIW, I agree with this viewpoint 100%. Thanks for stating it so clearly.
Domain names are nothing like programming languages.
That said, that would be the a very good step towards ameliorating the confusables issue in IDNs.
Sadly, no. There are actually huge swaths of Unicode disallowed in IDN [0]... although not everyone is aware of this and/or implements this check. Which is a shame, I'd really like to have e.g. "🃏-vd.name" for a homepage.
[0] https://www.iana.org/assignments/idna-tables-12.0.0/idna-tab...
It's a well-known semi-posh dish [1] (and yes it's an open sandwich since in Sweden that is the default), as well as a fun word since it contains all of our three national characters (åäö) at the same time, thus popular among programmers when dealing with i18n etc.
[1]: https://sv.wikipedia.org/wiki/R%C3%A4ksm%C3%B6rg%C3%A5s
Also, no I don't this is a good way to provoke a flame war with Swedes. I can barely register an emotion after seeing this.
Punycode is a bad compromise but nothing better is available. At this point entire TLDs now rely on it, it cannot be phased out any more. Hopefully this will be a lesson next time someone invents a protocol (though I doubt it will when I read through the comments of some anglophone commenters here).
There are so many examples of this:
- DNS (for a lot of other reasons than domain names)
- IPv4
- TCP (see the motivation for QUIC)
- a *lot* of incredibly obscure text encodings
- IBM 3270 control codes
- Sometimes even just ancient code no one understands anymore (search for the source code of Plan 9's `troff` for an example)
- …and so on and so forth
…and if someone wants to tackle such a situation, usually the outcome is either: - build another hack to add to the pile of hacks we won't be able to get rid of
- XKCD 927 (n+1 competing standards, because the old one will live on for ever anyway)
- face second system syndrome (resulting in an overcomplicated mess)
- get no traction because not enough people will switch to your new and shiny solution (due to the other calcified things in the stack or new flaws with the new solution)
So, my question: How should our civilization deal with this?Punycode is complexity, for sure, and it could perhaps have been avoided by just decreeing that non-ASCII labels are UTF-8 on the wire in DNS. This was considered, and you can imagine the lengthy threads that that topic produced then and still occasionally produces now at the IETF!
Bruce Schneier, more than 20 years ago (a short read, I highly recommend it):
Security Risks of Unicode
I don’t know if anyone has considered the security implications of this.
Unicode is just too complex to ever be secure.
https://www.schneier.com/crypto-gram/archives/2000/0715.html
What is the alternative here? Not supporting non-ascii chars?
E.g. for Finland .fi could do fine with just allowing ASCII + ä, ö, å.
But I have a strong bad feeling what's here already is going to remain here for the foreseaable future and the best we're gonna get is better handling rather than large changes to unicode itself.
In that case, you won’t have roundtrip conversions with any legacy encoding beyond Latin. For example, every legacy Cyrillic encoding treats the Latin A and the Cyrillic А as different letters. (Lest you try some sort of context-sensitive transform, both are in use as single-letter word: a French verb form and a Russian conjunction, respectively.)
Even if we could deal with that, the Cyrillic letter that is written as д (pronounced [d]) when used in printed Russian is written exactly like a Latin single-storey g when used in Bulgarian (never as a double-storey one). This is a problem for your idea even in vacuum, but every legacy encoding also treats this as a mere font difference.
As a further example, a and ɑ denote different sounds in the IPA (present simultaneously in French), yet there are fonts where the former looks like the latter (this is frequent in italics, but e.g. regular Futura and Andika do that as well). Those fonts are unsuitable for IPA, of course, but that shouldn’t mean a universal encoding must be unsuitable for it as well. Then there’s the Greek α, which is definitely not an a but kind of like an ɑ.
Do you want to distinguish the German Eszett ß (sometimes has a small protrusion on the left due to its origin as a ligature of long S + S/Z) and the Greek β (never has one)? The Greek τ, the Cyrillic т, the Latin m (used as the standard shape for т in Bulgarian), and the Latin(!) m-overbar (used as the standard shape for т in Serbian)?
I guess what I’m getting at is that the equivalence classes under “some language’s writing tradition has a glyph for C1 that is very similar to some other language’s writing tradition glyph for C2” end up much larger than anyone would ever want for general text-on-computers use.
Also. Perhaps domain names should be validated against the Unicode confusables list, forbidding confusable collisions.
Also, gonna bet that there are still confusables within character set constraints. A trivial example, the unicode confusable set considers "rn" and "m" to be confusable (in certain fonts and sizes). But I bet that's even more so in the broader european set.
Also it's a bit sad to be unable to, say, mix math symbols and english...
It seems to me that using the confusable set would be stricter (and safer) but also more flexible too...
Look closely, and you'll see it.
Your solution helps. But the better solution being pushed for, which is frankly the correct one, is "out-of-band" certification of domain names. Whether it's a little padlock next to the URL, or locking down the browser completely on non-certified domains. This will of course require more infrastructure, human and technical.
And, ugh, we'll get moral leakage.
Your advice is good. It's helpful if you don't fall victim to feeling too secure.
But, for the debate as a whole, I'm exhausted reading through everything like a lawyer. And dealing with it as a crisis instead of a meeting of minds.
On both my computer and phone, your "l" vs "1" swap is obvious even at a quick glance. But "lame" and "lаmе" look exactly the same in almost every font.
$ curl https://google.com/.curl.se
curl: (6) Could not resolve host: google.xn--com-qt0a.curl.se
Get it? That `/` after `google.com` isn't an ASCII `/`, and the TLD (se) doesn't have anything to do with this.OK, I guess then that IANA should disallow all TLD registries from registering such names.
Still, if the craziness comes from Unicode, one could be forgiven for wishing IDN had never happened and that Unicode in domainnames was not a thing.