A safer approach would be to only ever show a user the characters they expect to see (and are familiar with), e.g. based on their language setting. Assuming that every language has a finite list of characters used in its written form such a whitelist approach should be possible and much better than playing whack-a-mole with a blacklist for "potentially confusable" characters.
Oof.
Take for example, both 糉 and 糭 are valid characters in Chinese. One is a variant of the other. Which one is "canonical" depends on who (i.e. which authority, of which there are many) you ask. And FWIW the language and regional settings don't necessarily give an answer to the canonical representation.
Those characters mean the same thing with or without the specks of dust.
So, what's your solution here?
To be fair, Unicode domains are inherently a huge mess. The thing is that we don't need more armchair experts dreaming up Euro-centric solutions.
But as someone said, tiny, valid differences are easy to miss anyway, and original URL attacks were replacing similar-looking ASCII graphemes (eg. l for 1), so this will all continue.
That's totally backwards. If the assumption is that users of language X will legitimately visit sites of language Y with sufficient frequency, then all that language-specific filtering makes no sense.
And seeing punycode in URL bar does not mean a site does not work, it's only a suboptimal experience.
Either one or both of the characters you mention are probably part of the script of the users language setting in chinese; if the character is then it should be rendered as unicode and if not as punycode. If the users language has this kind of ambiguity then they are the only ones to judge if the domain name is correct or not, but at least they are familiar with the language and do not see characters they might have never encountered before and/or need to deal with an ambiguity that they shouldn't even have to expect to begin with.
The idea I proposed would still protect someone with a chinese language setting from being tricked by e.g. a cyrillic character in an otherwise ASCII domain name. I don't see how that is euro-centric (apart from ASCII being inherently english-centric), it is an overall improvement over the status quo no matter where you live and what language you speak.
Ah, funny. HN renders the ķ as punycode in urls. In the interest of sharing negative results, I leave this here.
Because there's a quick 302 redirect from "ķeepass.info" to "keepass.info" :
Chrome F12 Dev Tools network trace: https://imgur.com/a/vrxjsUV
Whether that redirect was there at the time of the Arstechnica article, I don't know.
EDIT ADD: around 12:57 UTC, the 302 redirect was changed to a Youtube video: https://imgur.com/a/TtLxafP
(Somebody is apparently having fun trolling the internet.)
ICANN lookup trivia says "ķeepass.info" domain was created 3 days ago:
Domain Information
Name: xn--eepass-vbb.info
Internationalized Domain Name: ķeepass.info
Registry Domain ID: a375f89abb384328a10460509f9f99f8-DONUTS
Domain Status:
clientTransferProhibited
addPeriod
Nameservers:
leia.ns.cloudflare.com
sevki.ns.cloudflare.com
Dates
Registry Expiration: 2024-10-16 10:21:45 UTC
Updated: 2023-10-19 11:40:19 UTC
Created: 2023-10-16 10:21:45 UTC