Show HN: UUIDs that are Shakespearean, grammatically correct sentences
github.com
github.com
For human-id I just went with groups of the 100 most common adjetives, nouns and verbs combined together in roughly the order they would appear in a sentence.
The output is nonsenese, of course, but I hope the combination of words understandable to anyone with a basic knowlegde of English might help.
I personally don't think the gramatically correct sentence helps here. To me,
"may-hold-come-foreign-low-white-cold-team-point-study-others-home-service-body-child"
and
"Cathleen d Dieball the Monolith of Alderson reflects Arly Arnie Keenan and 18 large ants"
Are equally poor representations of a UUID. Typically "human IDs" are useful for quick recognition or transferring over "voice" which neither of these approaches work well with (I'm biased, but IMO the first one is better for transferring over voice due to simpler words).
Human IDs shine when you have a comparitively tiny set of "things", like nodes in a cluster or Docker containers, and humans are heavily involved with accessing them. In this case a UUID is overkill but you want the UUID like gaurantees of universal uniqueness combined with more human recognizable identifiers.
An ID like "appear-hard-idea-case" is good because it nets you a high number of unique names but is also short enough to remember and speak aloud with a high probability that someone listening to you can type it out, word for word, with no issue.
tl;dr - never be in a position to need to have to "speak" or memorize a UUID.
From the page:
> Jacquette Brandtr Johm the Pectus of Barnsdall doubted Glenn Gay Gregg and 12 noisy stoats
As a non-native English speaker Brandtr doesn't look like an English word, Johm looks like a typo, Jacquette looks French (and contains a cluster of four letters for a single k sound) and I wasn't sure what a stoat was (my first guess was a type of goat).
I definitely find orf's version a lot more practical.
A stoat is one of these: https://en.wikipedia.org/wiki/Stoat
I just tried one now, it gave me "Kerianne Granny Dorise the Pulchritude of Belknap preserved Anissa Gareth Elery and 26 dead bees"
What are the chances of someone spelling "Pulchritude" correctly ? Especially a non-native/non-fluent person ?
And will an American spell "Dorise" as "Dorize" if you're telling them that word on the phone ?
And the case-sensitivity of the whole thing ?
Just give me old fashioned UUID's any way of the week ! Case-insensitive and no concerns about spelling.
I'd pronounce it as "Doris", maybe adding "-with-an-e" but a French person would pronounce it "Dohreeze" according to the internets.
Dunno if that would make an American lean toward "Dorise" rather than "Dorize" though.
Now that I know it means "to become Doric in manner or style", I do not know how I could have missed this obvious meaning :)
Depends very much on whether they studied Latin.
I think a lot of people would see "Dorise" in the abpve phrase and write "Doris".
if the UUID can be set arbitrarily then that's just another label.
That's pretty hilarious
This is what happens when engineers try to be playful.
I love projects like these but it's a bit of a stretch to argue that this one is practically useful.
http://www.pangloss.com/seidel/Shaker/
It only encodes ~18 bits of information per insult, but it’s a good starting point.
I can see how it might speed up writing say, SQL queries if you don't have to look up the UUID but not sure that this warrants the effort of committing this to memory.
For each UUID I 1) truncated it, 2) gave it a background colour based on the UUID itself, and 3) prefixed it with an icon for the object type.
The overall effect was actually pretty good. One could easily scan the list for the same UUIDs based on the colour, and the icon added clarity as to what was being referred to
For example, instead of invoice #306889086579, use an ID like R2020-06-14951.
Them being guessable shouldn't be an issue if you have proper access controls.
...which means that not being random opens you up to potential bugs and security issues...
I think I changed my own mind while typing this and I agree with you.
One of the benefits of UUIDs is that you don't have to check for duplicate IDs since the chance of generating the same UUID twice is effectively 0. If you use a system with less randomness then you do have to implement a duplicate check. Which isn't practical when you're working at large scale.
It's not memorizing them that's the goal. It's communicating them to another person.
IPFS and some of its predecessors have had to struggle with this problem. If you want an identifier that relates to the contents of the file, then you have an identifier that can't be transcribed easily.
The problem is though that the density doesn't go up very fast with additional words. Another poster pointed to his solution that only involves 300 words, and to represent a UUID takes 15 words.
If you made the dictionary 600 words it still takes 14 words to represent a UUID. :/
It's built on Cloudflare Workers so it scales fine and doesn't cost me anything since I use Workers for other projects anyway.
cat /proc/sys/kernel/random/uuid uuid -r
See https://linux.die.net/man/1/uuidMaybe your tool is different from the one installed on my machine because when I use the `-r` option I just get garbage back. This leads me to believe it's returning the uuid as raw bytes when I pass that option.
$ uuid
97af2898-cb54-11ea-8d75-176c10241ffd
$ uuid -v4
0b8d75f5-a148-466c-b1b5-9d9d1022c327
$ uuid -r
����T�����.
[1] https://en.wikipedia.org/wiki/Universally_unique_identifier#...This is considered shakespearean? Cool idea, but i think the implementation needs some work.
Example from Thepsis:
ALL. Goodness gracious How audacious Earth is spacious Why come here? Our impeding Their proceeding Were good breeding That is clear.
DIA. Jupiter, hear my plea. Upon the mount if they light. There'll be an end of me. I won't be seen by daylight.
AP. Tartarus is the place These scoundrels you should send to-- Should they behold my face. My influence there's an end to.
If that's not your thing try Tennessee Williams or Arthur Miller. Gilbert and Sullivan are all public domain so that's nice.
You could also go dramatically the other way and use flytings for your corpus: https://en.m.wikipedia.org/wiki/Flyting
There's probably clever ways to get the right entropy from their body of work while not turning it into word salad, best of luck.
And if also give you something to read in your head when you see it.
"Oh, you mean the 18 ants object?"
The point of the UUID is that it’s got enough entropy to be unique. If you start reducing it to smaller slices of its entropy, it isn’t a UUID anymore.
I think projects like this that try to make high entropy things appear more human-manageable do more harm than good.
> If you start reducing it to smaller slices of its entropy
Which it doesn't, as has been explained in the comments already several times.
Doesn't change the fact that it allow easier recognition, but you need to double-check it. (Not sure why I need to explicitly state that)
There should never be a situation in which the disclosure of a UUID breaks the trust or security model of your application your design is wrong.
A slightly different take on the same sort of problem.
Is it, though?
Plenty of folks memorise 16-digit credit card numbers - I've known retail employees who can recite those back after reading them just once.
Back when I was a sysadmin, I taught myself to type 25-digit Windows product keys from memory.
32 digits doesn't seem an unreasonable stretch, given time and practice.
Someone memorized 70,000 digits of pi: https://www.guinnessworldrecords.com/world-records/most-pi-p...
If you use a combined uuid (partially random, partially by date-time), it could be easier to remember 2020-07-15-ch99izy332 or similar... then you can take the date portion which is easier to segment (for some) and a shorter string of the subset of characters used for VIN or similar... all lowercase, and making 1 and l the same as well as 0/O, 5/S, and a few others.
Just some thoughts on this. For that matter, you could use emoji for some of those glyphs.
Algorithm wizards, can I get error correction if I dedicate a character for that?
The entire point of UUIDs is I can quickly generate them knowing that they will be universally unique, I don’t need to check for their existence anywhere.
This dramatically increases the likelihood of collision to the point I can almost certainly guarantee that they won’t be unique in any non-trivial context.
This still has all eight bytes, from the generated UUID, and looks like it can be converted back into a UUID.
edit: It seems the shakespearean uuids are longer than uuids so I may be conpletely wrong
However, an online demo will be nice to have.
UUID = --> universally unique <-- identifier
Why reduce the entropy just to make it look pretty ?
As for the people who say oh, but I can't remember/recognise "e0e93156-c68b-493d-bf31-19048db7dd9e"...
Well sure, but git invented that wheel before you. Just use the last 11 characters for local discussions/notes/whatever.
Finally, what about people who are non-native/fluent in English ?
"e0e93156-c68b-493d-bf31-19048db7dd9e" flows accross borders and languages
"Romeo Romeo Where For Art Thou Romeo" could easily be meaningless gibberish for a non-native/fluent English speaker, and also opens up totally un-necessary issues of pronunciation and spelling.
There is also the possibility certain words might mean something different in another language. Such as the famous Colgate in Spanish which means "go hang yourself" !
You assume a loss of entropy, why? The usage section describes a way to get a "readable UUID" by passing an existing, regular UUID. The reverse operation doesn't seem to be implemented yet though.
const id = require('uuid-readable')
const uuid = require('uuid')
var buf = new Array()
uuid.v4({}, buf)
console.log(id(buf))But man, the way you asked it. Assumed it was not possible even after the author replies saying they can mathematically prove it.
Why ask the question if you're just going to stick your head in the ground?
Downvoted.