Show HN: Base24 binary-to-text encoding for humans
kuon.ch
kuon.ch
> B C D F G H J K M P Q R T V W X Y 2 3 4 6 7 8 9
they were 115 bits encoded in 24 characters
see also human-oriented base32 encoding:
https://philzimmermann.com/docs/human-oriented-base-32-encod...
which includes this nice trick:
> We have permuted the alphabet to make the more commonly occuring characters also be those that we think are easier to read, write, speak, and remember.
edit: to add, an interesting human-readable and memorable base52 alphabet that I've never found a use for is to use playing cards
I encountered something similar to this a while ago when watching a multiplayer mod of Ocarina of Time[0]. They use a string of inventory item symbols to denote the identity of the server to connect to. “Hook shot, hook shot, master sword, deku nut” is a whole lot easier to remember than a long string of ascii.
I guess it’d be something like a Base62 encoding.
Basically they removed vowels (except for y, if it counts as one) as non-vowels often include a vowel in their sound. A fact reinforced while teaching my toddler daughter letters, words, and numbers. On top of that, they removed l/1(/i, and also o/0), m/n, s/5, z. Not sure why they removed z. Perhaps because of 2?
I'm not sure this is universal either, sound-wise. I suppose it does count for English. Because 7 ("zeven") and 9 ("negen") in Dutch get confused when spoken, some people say "zeuven" instead of "zeven".
My guess would be because of C, since Z can be pronounced as "zee"
https://github.com/bitcoin/bips/blob/master/bip-0173.mediawi...
> Why not use an existing character set like RFC3548 or z-base-32? The character set is chosen to minimize ambiguity according to this visual similarity data, and the ordering is chosen to minimize the number of pairs of similar characters (according to the same data) that differ in more than 1 bit. As the checksum is chosen to maximize detection capabilities for low numbers of bit errors, this choice improves its performance under some error models.
The most famous example is Schneier's Solitaire [1]. It was a common encoding I liked to play with in HS classes even before reading Cryptonomicon. I still think about it sometimes when I read through a Duplicate Bridge story in that long syndicated newspaper column. (One of these days I will actually learn Bridge, maybe.)
For syllables, I use: syllables: %w[ ba be bi bo bu ca ce ci co cu da de di do du fa fe fi fo fu ga ge gi go gu ha he hi ho hu ja je ji jo ju ka ke ki ko ku la le li lo lu ma me mi mo mu na ne ni no nu pa pe pi po pu ra re ri ro ru sa se si so su ta te ti to tu va ve vi vo vu wa we wi wo wu xa xe xi xo xu ya ye yi yo yu za ze zi zo zu ],
I can dump an implementation somewhere if people are really curious
Try different languages. ra re ri ro ru is my favorite little run of most of them.
Also it's always amused me how dog noises are onomonopoeitized in GA English as "bark" or "woof", when dogs lack lips to make a labial plosive, and their tongues can't really form proper postalveolar approximants or velar stops. I think it has to do with how we hear the third formant.
I guess I'd transcribe it like...
/ɚa◌˞'/
It's almost like "rorch" but more glottal less velar.
https://en.m.wikipedia.org/wiki/Voiced_alveolar_and_postalve...
Interested.
I've put together a few encoding libraries for fun when I get bored. (base16, morse, etc.)
This one looks fun, particularly because it _might_ be possible to serialise it to sound and back, if I put in a little bit of effort, which is something I've done [0] once or twice.
So 0 - O, 1 - A, 2 -B... (with S for 6, and N for 9).
Then 00 -> Double O holding a pistol
OA - Your friend Oliver Anderson doing whatever Oliver Anderson does
Etc.
Then 0100 becomes your friend Oliver Anderson holding a pistol in your imagination and that’s easier to remember and you can make stories with your characters to remember phone numbers etc.
Once I started this comment I realized it may be just a tad bit more involved, but I wonder if you could combine the two and have characters for every number from 0000 - 9999.
Consonants [(None), k, s, t, n, h, m, y, r, w], followed by,
Vowels [a, i, u, e, o] forming 5x10 matrix,
+ semi-voiced ゜(p replaces h) and voiced ゛(g, z, d, b replaces k, a, t, h)signs,
+ silent “nn”,
- wi wu we.
(aka NES Dragon Quest spell of resurrection)
https://thedailywtf.com/articles/The-Automated-Curse-Generat...
Not OP, but me, personally, I don't care about accidental obscenity. It is accidental, after all.
Then again, I live in Germany, and we don't censor swear words on TV either, so this is likely a cultural thing.
Implements two dialects. An original one compatible with the inspiration, and used in some of our earlier product. And the replacement (syllables above) which is alphabetically sortable.
You could (almost) easily replace every symbol with a single unicode rune from an abugida like katakana/hiragana (you'd need to pull from several langs as japanese famously lacks distinction between La-li-lu-le-lo and Ra-ri-ru-re-ro (らりるれろ) but there's no reason why you couldn't encode one-rune-per-phoneme.
> The characters that are used in Open Location Codes were chosen by computing all possible 20 character combinations from 0-9A-Z and scoring them on how well they spell 10,000 words from over 30 languages. This was to avoid, as far as possible, Open Location Codes being generated that included recognisable words. The selected 20 character set is made up of "23456789CFGHJMPQRVWX". [2]
[1]: https://plus.codes [2]: https://github.com/google/open-location-code/blob/master/doc...
I've read this a dozen times. Isn't OP saying that their character list includes G and 6, which are _not_ present in that list?
Update: It appears to be a typo in the article. Here's the real alphabet (N replaced by G and L replaced by 6): ZAC2B3EF4GH5TK67P8RS9WXY
https://github.com/kuon/java-base24/blob/0c25905414f1598a0ed...
Sorry about that.
P R
2 Z
8 B
look similar, depending on the font
It would be better to include some lower case characters which have more visual variability than trying to obsess over an arbitrary, inflexible stylistic "design."
If it's ambiguous, you could accept either and transform it to the correct value (implicitly, or as entered, or whenever makes sense. your users don't ever have to know). Or if you can't do that / the differences matter, do something like 1password does with chars and letters: show them differently https://www.dropbox.com/s/a29g2uiggqujzjl/screen%20shot%2020...
That’s missing the point. You can show them differently, but the point of keys / recovery codes is that they’ll be stored somewhere and later re-entered. Users could store them in any program (including writing them down or printing them out), you can’t control how they are displayed over there. Then when they need to use them, there’s a chance the ambiguous characters can’t be easily discerned.
Or just try all combinations, unless they entered o0o0o0o0o0o0o0o0o0o0o0 you're probably only going to have to try a small handful.
However, I strongly dislike the arbitrary mapping between character values and base-24 digits. There is a strong reason for using the order 2345679ABCEFGHKRSTWXYZ, which is that now encoded values compare the same as the original binary values. I did appreciate the 0x00000000 == ZZZZZZZ equivalence, but consistent ordering is just way more important IMO. Also 2222222 looks a lot like ZZZZZZZ. Just saying.
Ordered, your snippet look like the alphabet with a few missing letters, and isn't searchable on google or anything. I really wanted the alphabet to stand out.
I don't think that it is important that it can be sorted, it is intended for randomly generated keys which by my experience, you won't be sorting.
127.0.0.1 lusab-babad
63.84.220.193 gutih-tugad
63.118.7.35 gutuk-bisog
[0] https://arxiv.org/html/0901.4016This 128-bits can also be represented in, let's say base-50K, by using five words chosen from a 50,000 word dictionary. If you also make "this", "This" and "THIS" separate, then you can get away with a 17K word dictionary. Depending on the language, if you use roots and then vary morphology based number, tense, etc., then the number of root words (and the choice you have in making them simple) can be reduced. Such "pass phrases" can be easier to remember, transcribe, etc. (Also you will get random, humorous, offensive, etc., phrases...)
How to build the dictionary? Well, in order to determine the most commonly used English words, I downloaded a bunch of free texts from Project Gutenberg, and did some simple filtering - nothing less than 5 letters, no duplication of singular + plural, etc...
A valuable lesson that I learned during this process is that when your corpus includes older english texts, you should always give your final list a visual once-over and apply some judicious manual filtering. I'm looking at you, "The Adventures of Tom Sawyer". (And, to a lesser extent, Moby Dick).
I have an application where I’m using a 32bit serial for the event someone has to read it to sales staff over the phone. I would have liked to use 64bit and encode some more details into the serial. This would satisfy that.
I like the idea of removing ambiguous chars. I have a Base64 system that prints where I and l use the same font (infuriating).
128bit can be then represented by just 13 characters, or even much less with modifiers.
Such padding mechanism should not be necessary, and the padding from standard base64 is also not necessary. If you remove the ==='s you can still unambiguously decode it (despite the error some tools will give). URL-safe base64 (RFC 4648 §5) does not require padding and can represent any data length.
The classics like "", "f", "fo", "foo", ..., "foobar" would suffice. If the encoding specifically works on numbers, put test vectors for those too.
This doesn't require the user to be technical.
Also, any string like this should have at least a check digit and ideally some ECC digits.
You can add a single check digit with good performance using the Damm algorithm: https://en.wikipedia.org/wiki/Damm_algorithm one of the external links on that article has a suitable quasigroup matrix for Z_24.
I know there's Mega or Giga that can describe how big the number is in decimal, but they can do better (represent bigger numbers) in the baseN method where N > 10. So will we shift to these methods?
I am sorry, and I updated it.