Understanding and avoiding visually ambiguous characters in IDs
gajus.com
gajus.com
While I'm not arguing that their decisions were wise nor that they shouldn't have been able to foresee and prevent the issues they caused you and your colleagues, I would add this one thought in response to the line quoted:
It's often either beneficial or at least considered beneficial to prevent business information leaking through serial numbers, the simplest example being that if you start labelling your products with 1, 2, 3.. and never deviate, then it's fairly easy to take a sample of not many serial numbers and estimate how high they go and therefore how many have been sold. Sometimes it can also be beneficial to make it harder to guess a valid serial number (eg it prevent customers from pretending to have a valid one to get a refund, or whatever).
Of course, even if you have these concerns and want to mitigate them, it doesn't prevent you from also taking steps to prevent difficulty reading the correct characters. If anything it should make them more aware of the potential issues you faced since it means someone is already actually thinking specifically about what system to use, as opposed to what likely happened in your case of someone spending 30 seconds going "we need serial numbers, using X digits means we'll never run out, job done".
Also known as the German Tank Problem[1].
I think only consonants and digits are used in device serial numbers.
An unfortunate example. That's TIED HITCH SLOE REIGN RULE MOW? With only two parity bits, you can't even be sure this decoding is invalid.
RFC 1751 [0], from which this example comes, doesn't envisage the encoding being used in oral communication. Instead, it makes codes easier the user to "read, remember, and type in".
For oral transmission among professionals, sticking to the 26 upper case letters and relying on the NATO alphabet for encoding is a reasonable choice. Getting codes from untrained users in a lossy oral environment is still an unsolved problem.
Typing something letter by letter in Latin when neither party is a native speaker of English is very much painful almost half the time it happens
My wife thought it was crazy the first time she heard me use it. Then she realized that they all understand it too.
Some people are going to sound like they’re saying Todd for Tide, and have you heard how Baltimore pronounces Iron?
> These require use of a keyed message-digest algorithm, MD5 [Riv92] […] while sufficiently strong […]
Heh!
> […] is hard for most people to read, remember, and type in.
Ok, go on…
> English words are significantly easier for people to both remember and type.
Most people don’t know English.. But that shouldn’t be a problem since the word list can be changed. Right?
> Because of the need for interoperability, it is undesirable to have different dictionaries for different languages.
Oh. Well the world already learned the 26 characters of the English alphabet so adding a few words is probably fine..
> char Wp[2048][4] = […]
Oh, well at least it’s common words suitable for English beginners?
> WAD, BESS, MERT…
Hold on, these words are tricky even for…
> ORR? AGEE EGAN HAAS!!
…Are you done?
> GAUL FLAM! DRAB!
One day while sick, I distracted myself from being sick by writing up a silly module to do arithmetic in arbitrary bases. And, because it was easy I stuck it on CPAN. https://metacpan.org/pod/Math::Fleximal is the module.
Of all of the silly things I'd done, I would have sworn that this is the one that should never generate a support request. But it did! Why? Well I'd included a demonstration of how to turn hexadecimal into an alphanumeric code. And someone had the bright idea of using the same thing to turn long numbers into readable codes!
My module worked, but I was still a bit flabbergasted that THIS wound up in production somewhere!!
It helps if you draw a horizontal bar on the 7 but many don't, so you can never really be sure if a 7 is in fact a 1 with the serif or vice versa.
I.e. for a given number of IDs, how many characters are needed in the 53 versus 22 encoding (people who are not good at math might assume it is more than twice as many).
https://is.mediadelivery.fi/img/468/a93c32e08dae4768869a4bda...
No chance of confusion. This seems to have prompted some to add the serif to their 1 for stylistic reasons or whatever, since it's still distinguishable from 7 with a bar.
But then again people following older or newer conventions drop the bar from their 7:
https://is.mediadelivery.fi/img/468/46827e3320294f89b12a9338...
This makes a singular 1 with sloppily drawn serif hard to distinguish from a 7 without horizontal bar unless you can also see how the same person draws the other digit in their style.
See the last example in this image:
https://upload.wikimedia.org/wikipedia/commons/thumb/e/ee/Ha...
Side note to OP and author, the Wikipedia page is pretty handy and has a lot of info:
https://en.wikipedia.org/wiki/Regional_handwriting_variation
It never gets confused with 1, but in America, people were confusing it with 9 (!!), so I had to stop writing it like that. Can't please everybody...
My handwriting has always been pretty sloppy. My 9s come out like your 7s when I don't close the loop properly (I start at the bottom).
People confuse my lowercase r's for n's all the time too for a similar reason. Either I loop a little too much or I drag down the overhang so it basically is an n.
so 'muricans mistook my German ones for sevens, all the time, and I had to force myself to write what looks like a pipe symbol vertical bar to me instead of my trusted one.
and to disambiguate, we cross the seven like a lower case eff or tee is crossed.
I'm English, and I can't honestly remember which country it was that I've lived in (I think France...) where there were a couple of numbers that even after living there for a year I still wasn't confident reading when hand-written on things like café menus. And I don't think I would have thought of that being a systemic issue rather than just blaming an individual's handwriting before I lived there, despite having taken over 100 trips to France before moving to live there for a year.
Here's a deep link to someone in Germany writing down what visually looks like "77.5 :7:7" but his narration says it's actually "11.5 :1:1"
But this thread reminds me of when I lived in Canada for a while (coming from France) and I did misread numbers very often, which was totally unexpected to me. Yes, 7s and 1s looks very different between Canada (and the US I guess) and France (and probably the rest of Europe).
I haven't had this problem with Belgium though I'm not surprised if the standard here had been chosen to be the same as in France.
I was just saying it was obvious to me and it even takes effort to see how they could be misinterpreted. But I know they can be.
I was born in Europe so I put a horizontal line midway through 7. But now I'm in Canada and nobody else does. It can be a really tiny angular difference between a 1 and a 7 for a lot of people! :)
But it did not mention the most similar-sounding pair "F" (Foxtrot) and "S" (Sierra), which are nearly indistinguishable.
While one could use the NATO/Aviation standard alphabet (Alpha, Bravo, Charlie, Delta...), unless you have a very specifically constrained customer base,it won't help much. Best to also avoid those combinations.
Definitely better to have a slightly longer ID_String and maximal ability to read and speak/hear the characters. It'll save FAR more time and aggravation.
The end-to-end transmission can get really bad when you combine several different filter stages, such as a speaker's mouth being injured or obscured, a narrow channel like telephone or radio, noise, and a listener's ear losing parts of the spectrum.
As the sound transmission gets worse, you can get more rhyming ambiguities. Effectively, the consonants are lost in a bad channel and only the vowels come through. In an American English accent, I think these are the groups corresponding to different vowel sounds: A/H/J/K, B/C/D/E/G/P/T/V/Z, I/Y, O, Q/U, F/L/M/N/S/X, R. "W" stands alone with multiple syllables.
Depending on the kind of transmission problem, these groups can start to split apart into smaller subgroups based on which of their sonic differences make it through to the listener.
My family name begins with a 'F' and, indeed, I can't count the number of times where people write a 'S' instead. I've got invoices with a 'S' instead of a 'F'!
When I reported this bug they said it was for convenience!
I'm not sure if this UX was built into the OS, or just part of the game I was playing (Mario + Rabbids Sparks of Hope).
This is a extremely simple idea, but especially with random passwords this helps a lot even if the font is already hyperlegible.
We are not talking about encoding information only in color (= bad idea), we are talking about encoding information that is already present additionally in the color. And if your app has accessability settings (it should) this would be a thing that you could switch on and off.
Whenever I'm creating a 2FA backup on a piece of paper, anxiety hits me every time I cross over certain characters, o/0, v/u, 5/S, etc. I've come to add some fanciness to how I write these characters for this exact reason.
On "Phonetic similarity", reminds me of how I chose my wifi password. I wanted a common word with multiple consonants that a 3rd grader could spell, so I could share the password with a single phrase and have it be unambiguous. Ended up choosing "vacation".
It’s mind bottling.
It's probably fine to just print it out, but for more sensitive items I definitely write it down by hand.
Or perhaps they could: https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...
> They’re long and gibberish, odds of an unnoticed error are high.
That's why you "whitelist" those you wrote down and re-used with success: a little checkbox, which when checked means "Successfully re-initialized an authenticator with this 2FA?", works wonder.
A "dot" underneath a character means it's a number (so I'm sure not to mistake '5' with 'S', for example).
My "paper 2FAs" then go to the bank, in a safe.
I've never ever lost a 2FA access code.
I just bake the whitelisting into every 2FA code I handwrite. Instead of scanning the QR into the phone and then writing down the backup, I just start by writing down the backup, and then input it manually from the note into my phone. Once successfully used, I know the handwritten 2FA code is valid.
> A "dot" underneath a character means it's a number (so I'm sure not to mistake '5' with 'S', for example).
That one's good, I'll start doing that from now on! I also found writing letters partially in cursive to help too.
> My "paper 2FAs" then go to the bank, in a safe.
Yep same, I got a bank SD box back in 2017 during my first crypto wave. Have found the $100/yr to be incredibly useful. More recently I've created a sort of "defense in depth" for my passwords/codes. Least important things are available a button click away on Bitwarden Chrome extension, more important things are non-cloud-synced google-authenticator on my phone with 2FA backup in bank SD box. Most important things (i.e. crypto private keys) are sharded into pieces and distributed amongst multiple SD boxes.
My convention is that I put a dot '.' below every digit (this solves the 5/S, 0O, 8/B etc. issues [the actually problematic ones shall depend on your handwriting]).
If I'm really unsure, I add the NATO/aviation alphabet [1]. There's a 'U', I'll write 'Uniform' (in diagonal, starting from the 'U').
It only requires some discipline. I've done that since more than ten years now, never lost a single 2FA code.
[1] nitpicking about the actual difference between the NATO and aviation codes can safely be send to /dev/null
Some of these are areas of best practices that, when done really well -- people may not even notice it. That's an unfortunate fact of life that comes up often -- where the attention to detail and sincerity that people bring to the table often gets lumped under "obviously it should be that way, nothing special to see or applaud here".
9qg6G8B2Z5SIl170O (ariel)
The name of the font is Arial, not Ariel. (No mermaids here, move along)
https://github.com/gajus/gajus-com/blob/main/src/blogPosts/2...
I fixed the typo though. Thanks!
Or you should do the opposite - use real dates/words in ID and your visual confusion almost disappears (though there is a bunch of ambiguity here as well in similar pronunciation, so also not perfect). Humans aren't robots, so shouldn't be forced to read meaningless list of random letters
(example of geospatial system of coordinates based on that is what3words)
https://patents.google.com/patent/US9883333B2/en
I think it’s also a good example of increasing computer dependency by ‘human centric’ design: I can quickly and manually sort through a bunch of packages with coordinates or pluscodes written on them with some sense of locality. What3Words is designed to give a sense of familiarity but require an API lookup for every single address.
Letters and numbers also translate directly in most languages, words don’t (take bow as an example. Is it when someone leans over, an archer’s weapon of choice, or a cutesy headpiece?), so the familiarity aspect is limited to people with a good grasp of English.
Its main feature is that it can be commercialized, unlike regular coordinate systems.
Front of a ship, duh.
Free speech is a right. Interopability should be a right. Any infringement of those rights better gave a damned good reason. It's profitable isn't a good reason.
Example, docs and links here: https://www.mclean.net.nz/cpan/couponcode/
The world has almost unanimously decided my name is now Lain.
Fortunately with names, there are no returns, but exchanges are accepted (with a low restocking fee) in perpetuity.
Recently I went through the process of changing my name legally, because I'd fallen into a bad habit of writing "Steve" when asked for my name on some documents, but then remembering my "official" name was "Steven" on others.
Having multiple IDs with different names, especially after moving to a new country, was just too much of a pain - for example my official residence permit name didn't match my passport name, which caused some fun at airports.
I have some other context dependent characters/letters.
I write small z like that in normal writing, but as a mathematical variable I write it as ƶ. (To disambiguate from 2.)
I write small t like † in normal writing, but as a mathematical variable I write it as t. (To disambiguate from + (plus).)
I write q like that in normal writing, but as a mathematical variable I write it with a stroke, which does not display on the iPhone, a ꝗ, a bit similar to a ɋ. (To disambiguate from a (ɑ).)
It’s all about disambiguation, and sometimes having different letter shapes for isolated characters.
Thus after normalization, '1lI' would become '111'. This allows you to add seven characters back to the author's code generation alphabet without re-introducing any ambiguity.
However, if you don't need them, I would remove them so that the user doesn't have to spend any time wondering which character it is. Even though you're processing them all after they type them and fixing them, the user has spent time and effort that they didn't need to, just picking which one it is.
IIRC, I chose to keep them when I did something like this, but I don't think I thought to accept the others and convert them automatically. That project is sunset now, so it's not an issue.
Tangent: All number started with 12 which in effect made them 10 digits. They worked together with a banking system and the bank folks thought 10 digits was not secure enough so they complied and added 12 in front of everything.
Delicious malicious compliance - I like it.
(There's also a “footnote” by Donald Knuth: https://www.tug.org/TUGboat/tb35-3/tb111knut-zero.pdf, and follow-up by Bigelow: https://tug.org/TUGboat/tb36-3/tb114bigelow.pdf)
I don't know. People tend to use the letter 'O' a lot. And people tend to use zero '0' a lot too.
Who gives a fuck about "Oh"? I mean, seriously, which percentage of articles, blog, PDFs, webpages, products etc. throughout the world have have 'O' and '0' that can be mistaken one for another? And which percentage have "Oh"?
When was the last time a user had to read a product ID over the phone and did misread big O / "Oh" for 0?
I don't even think there was a last time, because nobody is using "Oh" in identifiers.
While, on the other hand, it's perfectly fine to use a slashed-zero for zero, to be sure nobody mistakes it for the letter 'O'.
So basically: your link and TFA aren't that related.
So, assuming (still not clear from your comment) that you do understand "oh" to mean the letter 'O', as intended, still your comment is surprising, because some of your own other comments talk about O/0, and the submitted post here too starts with that very example:
> What are visually ambiguous characters?
> O / 0 - The letter O and the number 0 can look very similar
So surely the article is relevant to (at least the first example of) the post? I admit it goes much deeper into just this one example, and only a bit into other examples like 1/l/I and 2/Z or 5/S, but still it's relevant and of value as a representative example I think.
The type face you linked is not optimized for humans.
> Its monospaced letters and numbers are slightly disproportionate to prevent easy modification and to improve machine readability.
It's a slightly different issue than what was described in the article (e.g it can't address the cases where IDs are written down).
In many cases these kinds of IDs are just an encoding of a ground-truth that is a big integer or a sequence of bytes, and that mean we don't have to use ASCII-character granularity, we can also use words.
True, that creates a certain cultural bias for wherever you get the words from, but it opens up new possibilities for error correction and detection, both by the computer and also by the humans transcribing things.
https://twitter.com/jonty/status/1570062564523917312
> the actual address should be "keen.lifted.fired" instead of "keen.listed.fired" and someone clearly misheard over the phone
That scoring/clustering process makes for interesting problems in their own right, especially if one throws accents into the mix.
I'll happily boycott that for-profit company which is masquerading as a public utility, but charging money and going after anyone who reverse engineers what words are what locations.
See also the comments in https://news.ycombinator.com/item?id=27058271
This is exactly the sort of thing that shouldn't be a private company, just like Lat/Lon coordinates and street addresses are effectively public domain, any suitable replacement for lat/lon should also be public domain.
If you do, you're not storing your bits as text to begin with.
I don't think words work well for codes that aren't meant to memorized. They make it harder to currate a unambiguous list since that list needs to be several orders of magnitude larger and the ambiguity can accent dependent. Of course, if memorization may be needed, then that is effort may be worthwhile.
Error detection with codes isn't hard, that's why checksums exist.
However, for the core purpose of the phonetic transmission, it seems needlessly verbose and cumbersome. The short wordlist combines with some fairly long component words to make the phonetic representation unnecessarily long. Additionally, I'm not super into some of the fairly obscure names and words included on that list. If I don't need memorability and hexadecimal atomicity, it doesn't seem worth using.
And we do, Bravo for B, Papa for P: https://en.wikipedia.org/wiki/NATO_phonetic_alphabet
Always use phonetic code if you're transcribing letters to someone, especially over phone/radio. It saves a lot of hassle on both sides.
If you don't remember the code, no big deal: For everyday situations, use any easily understood word. Like Apple for A.
And avoiding vowels can help avoid offensive words within a generated code:
FUKFUK9 - https://www.replacements.com/china-fukagawa-fuk9/c/27446
KUNT1 - https://id.made-in-china.com/co_gzberlin/product_Power-Steer...
base32 removes the I,O,U but other words with A,E need to be avoided too - no vowels helps avoid words in English.
Showing cl and d can be hard to discern clifference.
-B, --ambiguous Don't use characters that could be confused by the user when printed, such as 'l' and '1', or '0' or 'O'. This reduces the number of possible passwords significantly, and as such reduces the quality of the passwords. It may be useful for users who have bad vision, but in general use of this option is not recommended.
>pwgen -B 32 oos9upoVieghuew7aeb3iev3jiequeiw acohthahpie7ae4aeboshahWiengieth yahW3qua3atheeP9jo4aiY3zeepoosh3 Noh4ooth4ohzeec4zug3ephoo7meich7 oozae9Eireix4Chaiboz9dofie4Xunof Mohj3uupee9ahngahh9on9sujee9ehae weimah9aiXeis3owaexei4uh3ibeecai PaeV7eeChaezahruNgeequoh7zok7thi eeJieyah4exiephaiPootei4dokoojoh fohhah3Eec3bah7aeR9iedah7Ve3ea7o vahs4eich4pheisoug9aiR3ohChoh7Ch eth9KaeLahdie7ahy9ohCiebohphuse9 ieye3udumaengai9ies7kae4geeque9T iesoh9eosohthoongaeroo4ehiishohY mee4ohjei4ohmika3taijei3Yaixosei ohWoo4eapid7miebee9pooKai3oofeis Eechook9quohp7se7ees9thaefahb9an aht3quooV4eiph9ap7aiw4wee7oi7eij ishep3weeh7Eero9ohdohth9MietooJ4 Kai9aich9Jee9Angeihee9eehei9esie toonaix4xe3Moob3zaic3Eesahs9ahy3 gaey9doozee7sei9quuPae3vohph4Huo ouYaephahcog3peiw7iecoo7eetheeph eeNgiezae7oongi7uena7eenaezuT7co tai9vuace9eV7Paih7ieN3Ahghiegh3v VaeteeMoobeixai9ingeyahYuzaipaht eeng7vei7pho4Ahpoa4kahgheethahz7 phas4theiThu4uqu7iCh3Aepha3shae3 ieRep3kaideeHeekiNgequieng9raeYo eegahsh9aizooshee9too9oojiox4Lei ovohcaePahM9thaebajuChoo3pipheej oowaimeiWahf4Neighoo3Eeyah3uvi4v vi4choiThei3eisohw4iP9huehohs4oe ukuchiethaquax3hieChouMahpooy4ee aegheeyeemeNeevehud9ohng3dai4jai eth3iedah9Tee3wohneisoo4aicuToos iecap7EeJ7raixiuseesiNou9ooT9fie ied3ooveingu7fu7dahdaaYe9tai7ien eijee7iKighaingaiChei7giemu4chi3 Thie3faih3ahshooRunohwoaghoh4Aev
https://philzimmermann.com/docs/human-oriented-base-32-encod...
Command line tool at https://github.com/tv42/zbase32
$ echo hello, world | zbase32-encode
pb1sa5dxfoo8q551pt1yw
$ entropy 16 | zbase32-encode
y64s31aq6cgjoko9fwbuasf4ce Term: ASCII digits
Example: 0123456789 U+0030..U+0039
Explanation/Description:
Commonly used with Latin, Greek, Cyrillic and many other scripts, including some non-European scripts. Used in alternation with native digits in scripts that have them. (Some scripts with native digits make only limited use of ASCII digits.) Infrequently used in many of the remaining scripts.
Synonyms: Western digits, Latin digits, European digits
Which then links on to: https://www.unicode.org/glossary/#european_digits> European Digits. Forms of decimal digits first used in Europe and now used worldwide. Historically, these digits were derived from the Arabic digits; they are sometimes called “Arabic numerals,” but this nomenclature leads to confusion with the real Arabic-Indic digits. Also called "Western digits" and "Latin digits." See Terminology for Digits for additional information on terminology related to digits.
I think anyone who has dealt with both Arabic numerals (as used in Europe) and Arabic numerals (as used in parts of the Arabian world) feels the naming is unfortunate. Arguably this is not the best place to bring that up, but I certainly stopped using "Arabic numerals" after working with some i18n code which supported both Arabic and Arabic numerals.
I heard an American making a joke that
"I have gg problems but European handwriting ain't 7 of them."
This takes the approach of allowing ambiguous characters by decoding them to the same value, and also considers the problem of accidental obscenities.
5-bit base-32 oi23456789 abcdefghkl mnpqrstuvw y
o = 0 i = 1 j, x and z removed.
I like that you can fit 6 characters in an 32-bit integer and still have to bits to spare... makes for compact usernames and network bandwidth.
1. IDs can be generated anywhere (client-side, server-side, etc.) and are still unique 2. IDs are ordered by time 3. IDs don't use L and O because those can be confused for other characters
I've found it very handy in my travels.
[1] https://github.com/stevesimmons/uuid7-csharp?tab=readme-ov-f...
https://github.com/bitcoin/bips/blob/master/bip-0173.mediawi...
But I use a alphabet with 32 characters: abcdefghikmnopqrstuvwxyz23456789
I prefer 32 characters, because that makes it possible to pack 5 random bytes into a token with 8 characters.
Since the whole point is the ability to convey a message in the physical world end with chalk or pencil or whatever – we needed to make sure that characters were unambiguous.
So there are no zeros or ‘o’ characters or ones or ‘l’ characters… I think there were one or two other rules that govern this but I can’t think of them right now…
[1] https://0x.co
My research focused solely on the .com domain name space, so our character set was limited.
Essentially A-Z, 0-9, and the - character, and domain names can not start with the dash character.
xxx_flown-moons-deary-flake
I wish the author would have said more about this. Why be wary?
[0] https://unicode.org/reports/tr39/And even on systems which do have these fonts, they may not always be exactly the same.
Nitpick, but isn't this polynomial to the members of the set?
This is why longer password are more efficient than complex passwords: to gain the same security effect of doubling the password length you would need to square the alphabet
We're comparing the growth rate of of two exponentials representing variable-length identifiers. We're not looking at a constant-length identifier (which is what you're doing with only looking at a^n). Notice the context of where exponential is used in the article: we are changing n from 5 to 8.
After that there's ECC. A few extra bytes for a reed-solomon code will fix a lot of issues.
They’re not even mentioned and don’t look like a thing else, except maybe each other in some typefaces.
This applies to usernames too! It's easy to phish if platforms render capital I and lowercase l the same
if someone's writing is incompetent tell them. if you can't then they ruined it for themselves by being shit at writing the number 7.
xxxxx-xxxxx-xxxxx-xxxxx
Instead of something like this: xxxxx-xx-xxxxx-xxx-xxxxx
Something could also be said about such scheme lacking the embedding of a checksum.Here's an IBAN (bank account number) in the EU (which thankfully are using a checksum as part of the account number):
LU29 0022 1712 5582 7000
^^
||
two checkdigits
Also some companies think they're "smart" because they pick numbers like this: LU29 002 0000 0001 8000
Repeating the same digit, usually a zero, a shitload of time ain't smart. It's fucking dumb.- We haven't solved this already? Who hasn't tried to read some code and couldn't tell O from 0 or l from 1, etc.?
- Aside from ambiguous characters you have to be aware of spelling and leet spelling. e.g., 53X, S3X, 5EX, etc.
- FFS stop with the 10+ character strings without spaces or hyphens. There's no reason for that.
- Not everyone has perfect vision. Ambiguous characters *and* less than perfect vision (often with not spaces / hyphens) is a mortal UX sin.
We've all been on the wrong end of these, and yet they are common enough - in 2024??!!? - that they need to be mentioned here.
I was excited see that the post is getting engagement. I saw it in 3 position. Then checked an hour later and it is nowhere to be seen.
I am assuming this is some sort of opportunistic algorithm at play that gives a chance to a post, but removes it if it is not performing, but curious if anyone has more details.