Identity Beyond Usernames
lord.io
lord.io
I'm a screen reader user and don't encounter those often. There's even some fun in those alt texts, i.e.
"screenshot of unicode character inspector revealing "epic" to actually be "Đ”ŃŃŃ". of course you, a screen reader user, aren't fooled."
Btw, most people would just label that "screenshot", which is extremely infuriating.
Does it actually prevent the large media elements from downloading? how?
Maybe one day computers will be able to describe to me what is in an image without a person having to type it.
As an aside, Chrome does the autolabeling (including OCR) only if there's no alt provided, so alts like "screenshot", "photo" etc. actually cause more harm than good.
It also presents a front-end solution (fuzzy-text matching for @'ing someone's display name in a Slack channel) but does nothing to address the ever-present back-end need to actually validate that a client is who they say they are.
There's a good reason you can't make your bank username đ„đđ đđđđ , and it's not "because developers are anglo-centric and intrinsically against all the wonderful individual expression and identity that Unicode could bring us."
Since the article mentioned Punycode, I actually think the current implementation is an elegant solution to a real problem. Of course we should enable folks from anywhere around the world to experience the Internet in their native language. That's the reason the Chrome algorithm to determine whether to show Punycode is so complex - it has to carefully balance the desire to show text the way it was intended while considering that it might be a malicious attempt to phish someone using characters outside their typical locale. In this case, the front-end solution of asking the user "which google.com from this list of two identical-looking google.com's is the one you want" just won't fly.
To the system, of course they're different - it's just looking at the bytes. Problem comes when a human interacts with the account. Maybe they see the email on the account and manually type it in to another system. Or records are printed as part of a lawsuit and has to be transcribed.
Allowing indistinguishable characters in unique tokens creates confusion at best. Broadly I agree that we should reduce our reliance on unique identifiers, and I don't have a good answer for the right approach, but certainly we can't abandon the concept of a unique identifier altogether: it predates computers and the character encoding mess we made, by a long time.
I like to think I have a keen eye and pay attention to subtle details. I don't think I see a difference between the two addresses. Your message seems to state they're different though. Did HN normalize the differences? Did my browser?
In a blaze of stupidity, I decided to to copy and paste into my terminal and pipe it through a hex dump [0].
* The character difference doesn't show up in the viewport in Firefox; of course, it's rendered.
* The character difference doesn't show up in the HTML editor; of course, it's showing the "raw text" and the "raw text" is valid unicode.
* It doesn't show up in the browser's network inspection of the response payload. That's the scariest part in Firefox IMO.
* It doesn't show up in my Terminal either. Why shouldn't it? We've fought long and hard for terminals to support Unicode.
Of course, then there's the fact that I decided to copy and paste from a browser into my terminal even though I already knew I shouldn't [1]. What else could be hidden in unicode? An entire bash script starting with `sudo`, perhaps? Websites can already inject their shitware into the clipboard with clipboard events [1].
[0]: https://knightoftheinter.net/img/hacker%20news%20id%20237743...
That you think homoglyphs would somehow look different in any of those steps of your post is peculiar to me. l and I are the same in some fonts and they are in the ascii set.
See the Han unification [0] effort.
There are characters in Japanese and Chinese which are similar, but written differently in each... And they ended up using the same codepoint in unicode for them and relying on different fonts.
So now I can't easily quote a japanese sentence in a chinese book without having to use two different fonts, which seems quite silly.
Worse yet, there were many common glyphs in use (especially for names) that were unified out of existence. There are literally people who couldn't type their names in unicode. There are a lot of works of text that can't be faithfully OCRd due to not having certain character variants that were unified away.
Okay, so why wasn't the cyrillic alphabet unified with the latin one even if the japanese and chinese ones were? Clearly Han unification did much more damage than cyrillic unification would have.
Well, the answer is sorta politics. ISO 8859 is what came before Unicode, as far as the unicode consortium is concerned. Since ISO 8859 encoded latin and cyrillic separately, that got carried over for "compatibility". Because the unicode consortium and ISO 8859 were both more western-centric, and CJK users had already dealt with things in a way that was standardized only over there, not in any western ISO standard, of course the unicode consortium would honor the existing ISO standard and ignore the CJK standards.
Which glyphs were âunified out of existenceâ? Isn't it more just a matter of using the appropriate typeface or a variation selector? In my limited experience, I've noticed some things being non-unified that I would expect to be unified, like æ„ and æ© (e.g. in æŁæ„ [zh-TW] vs æŁæ© [ja-JP]); far as I can tell there isn't any semantic difference between these characters, æ°ćé« just added a stroke to make one of the radicals more consistent; I guess the idea was that Japanese users mix æ°ćé« and èćé« in text, and JIS character sets had separate codepoints for each from the beginning. As somebody who is learning both Taiwanese Mandarin and Japanese, I don't think it's common for unified characters to differ enough that they would be unrecognizable. I think that no matter how the Unicode Consortium approached this, they would need to set some limits on what gets a separate codepoint; the question is not whether or not to do the Han Unification, the question is when a glyph is actually its own character.
As for selecting specific glyphs of a character, if you have a font that even has the variant you're looking for, there is https://en.wikipedia.org/wiki/Variation_Selectors_Supplement
But the problem is, there are many characters very similar to each other, with a difference of one stroke already. ä» and 什 for example, are both simplified Chinese, but with completely unrelated meaning.
So, when you see a character that looks familiar, how can you tell whether you are looking at a character that you don't know or a character you know but rendered in a different language?
You canât code your way out of a context trap.
This isn't the best example of meaning, as neither one of them really means anything standing alone like that.
Let's talk about çŽ. Is that a japanese character? A chinese one? Well, let's look it up in a japanese dictionary [0] and then a chinese one [1].
You should see that the results look different. They're two different characters drawn in two different ways. However, in this hacker news comment, there's no way for me to indicate to use one font for one, and one font for the other. I can't say "The japanese glyph çŽ is the same unicode codepoint as the chinese glyph çŽ even though they render differently. They were unified". I can't make them render correctly as japanese and chinese respectively. For my computer, it _only_ renders in the chinese variant (without the extra stroke on the left) on hacker news. Like most sites, there's no way to indicate in the text input which language that portion of my text is. Unlike with every western script, if I don't indicate the language correctly, it will be actively rendered wrong and difficult for a reader to understand.
To draw an analogy from another hacker news comment, this would be like the unicode consortium saying that 'colour' always renders as 'color', and you just have to switch fonts for it to look like 'colour' [3].
Okay, so that's why unifying things at all is silly and causes trouble.
As for characters that were unified out of existence: unfortunately, examples of those are hard to give. There are various names that have stylistic choices or use unusual characters which can no longer be rendered "correctly". Arguably, that could be seen as akin to the fact that if you style your name calligraphically in the western world, unicode doesn't help you replicate that flair.
[0]: https://jisho.org/search/%E7%9B%B4%20%23kanji
[1]: https://www.mdbg.net/chinese/dictionary?page=worddict&wdrst=...
Sure you can, the Chinese one is ïȘš and the Japanese one is đŻ„. They are still the same character though, and it's meant the same thing the whole time. The Japanese got it earlier, so they form it in a way that would be recognizable to the scribes of the Zhou dynasty, and possibly the Shang dynasty. [0]
Part of how you can âtellâ it's the same character, is that the cousin variations are all used in precisely the same way. Many compounds formed with it are shared directly between Japanese and Chinese. [1] [2] [3] It's related right down to in some compounds being interchangeable with ćȘ, the latter being a less dated form of it (i.e. çŽäž vs. ćȘäž in Japanese). And beyond all of this, compounds shared between languages have been written in both orthographies and meant precisely the same thing for a very long time.
I think considering these variations to be separate characters makes about as much sense as considering the s in âstopâ to be different in French because it's pronounced slightly differently, and because French penmanship is different from British/American penmanship (i.e. sometimes French people lift the pen while forming a lowercase s [IIRC]).
[0]: http://xiaoxue.iis.sinica.edu.tw/yanbian?char=çŽ
[1]: https://en.wiktionary.org/wiki/çŽèš
While it might represent the same glyph, it certainly isn't the same sequence of bytes. I think the real failure is that's not made clear.
How would you solve homograph/glyph attacks though? One idea is yet another encoding where there are no homoglyphs, only whitelisted diacritic sequences, and there aren't more than one way to assemble the same character ("Ăł" vs "o"+"ÂŽ". So tough potatoes for Cyrillic "o", it's forced to use its nearest equivalent: 0x6f Latin "o" in the ascii set.
First, modern operating systems (should?) already provide APIs to canonicalize UTF.
Second, perhaps an additional API needs to be created which suggests similarities between characters intended for use by an intelligence (artificial or otherwise...).
pbpaste | xxd
Or your platform equivalent of pbpaste. If this isnât totally safe and could have side effects, your platformâs clipboard implementation is problematic.https://en.wikipedia.org/wiki/Email_address#Local-part
You can probably include Unicode in the "Friendly From" header which is what is shown in most email clients though.
If I log into AWS using one account (the one I've had for Amazon.com for more than a decade), I get the console with no resources in it. If I log in with the same email but a different password, I see all of my resources. Absolutely insane.
Could you effectively DOS an account by creating thousands of shadow accounts with different passwords?
The login handler is only going to try to bcrypt so many times before timing out.
You made me think of an interesting hack. We could pass public messages or make public statements that are forever ungoogleable. I could write: âHey, jĐŸsh, you can bring up the secret menu in Đ”ŃŃŃ games by clicking six times on the blue icon.â But youâd never find this message again by searching the Internet for josh or epic spelt the normal way.
By the way, a quick way to perceive the lookalike Unicode above is to paste the words with an appended .com into the address bar of your browser. If you paste jĐŸsh.com or Đ”ŃŃŃ.com, youâll see punycode.
Mailing addresses have the same problems to some extent, but at least most mail systems should have some concept of forwarding addresses by now. Email is never going to.
We have abstracted the notion of passwords as âa collection of authenticators defined by the userâ, no reason your account canât have a collection of identifiers as well.
The problem is that in practice, control of the email address on record is sufficient to take over many modern accounts. Bank accounts likely shouldn't fall into this category, but most web sites don't have bank account level of security.
Hotmail created a hornet's nest of problems when they started recycling email addresses after 6 (or 18?) months of inactivity. It allows a quick "password reset", then the account is now owned by whomever controls the email address. Effectively, the identity is hijacked because control of the email address was most of the authentication mechanism. Queue the spam messages and fraud/phishing.
Also developing/managing a customer service tool which allows them to decipher these "takeover" events and ensure the person contacting them used to own the account is difficult and sometimes not possible.
Just as long as you aren't copying those bank credentials into your clipboard :)
1. Push. Add a new item to the clipboard.
2. For apps, with an opt in permission. Delete items added by the app. Good for password managers.
Also, it's basically impossible to use a password manager without using the clipboard. At the one I use will automatically remove the items it sets after a timeout I can configure.
In my case I have zx2c4's pass and a Firefox extension named PassFF, so the extension just gets the usernames, passwords and anything else via IPC.
It does have a clipboard mode with an expiry timer, but I very rarely need that for anything.
Firefox's built-in password manager similarly doesn't use the clipboard, Google's doesn't, I believe Apple's doesn't.
It's surprising but providing an extra degree of freedom makes the system more robust. Username phishing is a real problem on Twitter because people expect unique names associated with each person they interact with but if usernames can not be assumed to be unique then using the name as a heuristic for identity is no longer a viable shortcut so people have to develop other ways of making sure they're talking to who they think they're talking to.
On a related note, keybase (keybase.io) proofs never made sense to me until I started thinking about how I would prove to people that I am indeed who I say I am. Keybase provides a cryptographic basis for trust, which is much better than what most social media systems currently support with their verification mechanisms. I personally trust cryptographic signatures over whatever verification mechanism Twitter is using to provide blue check marks to verified accounts.
I disagree that this is the only solution. The OP comes very close to discovering an alternative one: don't allow users the power to select their own arbitrary username.
When a user makes an account on your site (and if your site actually needs publicly-displayed usernames in the first place), give them the option of, say, ten usernames generated by the site itself; these don't need to be numeric codes like WeChat does, they can be pronounceable phrases in the same manner that Gfycat generates URLs. Everyone will end up with names like "questionable wet aplomado falcon", "each unlawful harlequin bug", "icy inferior iceland gull", which honestly isn't any worse than the average name I encounter on Reddit or elsewhere.
The biggest downside(?) of this approach is that it makes it harder for someone to build their personal brand, but the older I get the more I think it's a bad idea to use a consistent nickname across different websites.
I tried this approach for a service awhile back, and the biggest hurdle was educating the users. Even determining what to call it was a huge challenge, as itâs clearly not a âusernameâ.
You may be overestimating the typical non-hackernews user. Itâs at least as challenging as getting users to use things like 2fa and/or recovery codes.
This is the same UI flow as the above, except that it's not based on anything that the user has already entered and it doesn't provide any way for the user to override the suggestions, which actually makes it simpler than the existing flow.
Pronounceable in what language?
Why should Vietnamese or Icelandic users have to use English user names? It is really necessary to enforce English hegemony at such a fundamental layer of the software?
For comparison, imagine if your randomly generated user name was one of the following 5 (these are all combination of three words):
1. sao-chep-nong-pha-vay
2. giu-phep-ca-ngua-bay-chu
3. hai-hung-ngon-luan-boi-phan
4. hon-ba-kinh-can
5. xuoi-van-ri-do-truoc
and that's what you needed to remember when you signed in and what other people had to quote or @mention when they reply to you.--
The opposite is actually true - there are many reasons to change your email, while changing your username on the vast majority of websites is simply unnecessary (e.g. anything without a social component, such as banks, shops, Healthcare providers etc.).
Not to mention, if you don't have a system for ensuring unique usernames, you also shift the burden of ensuring identity to your other users, who now have to understand how to distinguish different users with the same name.
Even hints that for directories, etc at the end.
In particular, that number is only four base-10 digits long, and I'm pretty sure there are more than 10,000 people in the world who have ever played Hearthstone.
In many fashions, the account is identified by a "snowflake", a 64-bit ID space that every user, every guild/"server", every channel, and every message has.
Its described here and is the only permanent identity on the account, since everything else can be changed. The identity used to @ is transient and can change at any time, but the client, the API, and the myraid of bots written for Discord know the real identity is the snowflake ID.
https://discord.com/developers/docs/reference#snowflakes
At what point is it no longer an identity and instead a database identifier? That's a good question.
On @-mention autocompleters in general (and :emoji-code: completers too, and GitHub/GitLab issue/PR #-reference completers), I honestly canât think of one that Iâm completely happy with. Every last one Iâve experienced harms the editing experience with surprising and inconsistent behaviour.
Itâs all surprisingly hard to get right, and very few people that implement seem to even try to actually get it right.
Perhaps I should perform a more detailed study, enumerating the problems clearly, and try to create one thatâs as close to flawless as is possible. (Iâm sceptical that flawless is actually possible with the tools given web tech, especially in Blink and I think WebKit which donât properly support the difference between before-end and after-end in selection in contenteditable, which matters more than you might think.)
In spite of what I'd personally have bet on a decade ago, I don't actually know many international users who complain about alphanumeric handles. People can/should still have Unicode-ish display names on top of handles, change those, and be searched by those. Twitter does this well for example.
But a human-readable, ASCII-friendly, user-chosen unique handle is the best thing I've seen so far for disambiguation on a social network.
* IANAL, etc, and I'm sure requirements and interpretations of GDPR vary, but I do know that my company invested quite a lot into systems to fix this @-mentioning issue, so at least some people thinks it requires this and are willing to spend probably millions of dollars of engineering time to comply with said requirement. Personally, I support this provision.
1. Sites should forbid dots in "normal" usernames. 2. Usernames with dots can only be registered by demonstrating ownership of the corresponding DNS domain.
Now you can have your favorite "brand" across all sites, nobody else can register it before you, etc.
This is obviously not very non-techie friendly at this time...