The UX of UUIDs
unkey.dev
unkey.dev
No matter what your identifiers look like, if you want them to be easily copyable you should add `user-select: all` to the element containing them.
If you do this, all of the text will be selected automatically when you click on the element.
https://developer.mozilla.org/en-US/docs/Web/CSS/user-select
Also more than 3 clicks starts taking a considerable about of time.
> Customer: "I'm having issues with an order"
> CS: "Can you give me the order number?"
> Customer: "Sure, it's zero two a as in apple five..."
This seems entirely reasonable considering the whole point of TFA is to move away from using UUIDs as unique identifiers for resources in a service.
“Asking the customer to copy/paste the UUID from a URL to send to Support via an email or chat.”
Rather than asking a customer to read it out to you.
b214a2bb-c3c1-48fb-8272-10a2808337f3, something, ...
b214a2bb-c3c1-48fb-8272-10a2828337f3, something else...
Not as good as double-click selects all the id, but at least id doesn't take too much time and I don't have to precisely go to one of the ends.
I'm grateful to be informed about select-all, though!
That being said, most people copy and run bash script straight off the internet, so clearly not worried about copying stuff they haven't read!
e.g. https://bun.sh
curl -fsSL https://bun.sh/install | bashThe most common complaint about "pipe to bash" I've seen is the possibility for the server's response to detect it's being piped to bash, and then execute malicious code. The suggested remedy is to first download the install script (and check it) then run it. -- This seems overblown to me, since if you think the server may be malicious, then downloading programs from that server also seems risky.
Criticising people for not reading bash scripts from install pages is weirder to me. -- It's possible that some software author would hide malware in the install script; but, then why wouldn't they just hide malware in the installed program itself.
I heed the risk with the reasoning that even a benevolent server may be compromised, and that detecting pipe to bash is a potential way for that to go unnoticed.
No! That's throwing the baby out with the bathwater! Removing all separators means rare-but-important manual tasks of transcription or comparison become terrible, since there are no clear chunks.
Instead use a different character which doesn't have the same problem, one that most software considers part of the same "word"... such as the classic underscore.
For most people, double-clicking on this 123_456_789 will select all 9 important numbers. (And maybe a trailing space, but that's a separate problem.)
Like the coconut in TF2.
Obviously I can't fully express it in text here, but try to imagine this as a coworker speaking to you: "Hey, write down this IP address. It's ten, seventy, one twentyyyyyTWO, five."
They didn't actually say "period" or even "dot", but I bet you'd type 10.70.122.5 .
having a clear seperator helps me say the numbers faster
I've never once heard anyone drop out the dots in an IP. Non technical users aren't confident enough to do anything but read it exactly as it appears (one zero dot seven zero dot...) and technical users who are generally experienced enough to know what an IP address is, know that the dots are meaningful.
If it's something like 56.7.23.231, I'm definitely going to disambiguate it by deliberately saying each one of of those three dots.
But if it's more like 192.168.0.1, I'm probably not going to bother with speaking any delimiters in conversation with another person who has at least reasonable familiarity with common IP networking layouts.
Bringing it back to the topic: UUIDs should not ever follow familiar content patterns (if they do, then that's an issue in and of itself), so I'm always going to speak the delimiters of a UUID -- whatever they consist of.
(If nothing else, doing so breaks up the pattern into human-digestible chunks -- which is probably the sole reason we have those delimiters in UUIDs to begin with.)
There also isn't one for "w", yet we get by with that as a letter.
Warning: Tangential rant ahead.
I'm teaching my toddler to read (Distar alphabet).
Even with the modified alphabet, it's a chore to "know" how to pronounce a letter.
'a' has at least 4 different pronunciations in words used by toddlers: apple, came, eat, bread.
All the vowels are like that, and even some consonants ('y' has at the very least: baby, yesterday, cycle, buy)
The only well-behaved letter in English is 'x': pronounced the same wherever you see it, as 'cks'[1].
[1] For toddlers, anyway. I doubt a 4-year old would be interested in LaTeX :-)
Well, we didn't cover sounding out of `ph` yet, and it isn't a toy he has, so thankfully it is not a word he uses.
There's dub.
Dub-dub-dub is pretty widely understood to mean www.
But I get that it's confusing with dashes.
True, but alas, on iOS I found that a double-tap selected only one of the digit groups.
In my view the touch interface UX is just as significant - perhaps even more so in recent years - given that the backdrop to many of these identifier format decisions is ensuring nontechnical end-user support, under time pressure, over possibly quite unreliable channels, goes as well as it can.
But look on the bright side, at least it didn't try to call the number
0000-0000-0000-0000-0000-0000-0000-0000
Another VERY nice feature is it hashes each set of 6 digits as you type, so if you transpose one, you immediately get feedback instead of "invalid key!" after typing the whole thing out.
nice, now i can dictionary attack any key, 6 digits at a time
At worst, the key would have some portion less entropy since there's a lot of bits used for checksums.
Regardless of how many dashes you have or how (ir)regularly they are spaced, to select the whole ID you must carefully click-drag-release around its boundaries, you can't just double-click anywhere in it to select.
0000v0000v0000v0000v0000v0000v0000v0000
- K-sortable: ensures good locality when used as an id in a database (e.x. https://github.com/jetify-com/typeid https://github.com/segmentio/ksuid )
- checksum: primarily useful when an id might be conveyed verbally (e.x. customer support) or transcribed (e.x. Bitcoin wallet backup, BIP-39)
But if I could go back 25 years and only give myself one bit of advice, it would be to use UUIDs as the primary key. Because in a different context to raw performance, it offers a lot of advantages.
While there are advantages in numerous areas, I'll focus on one for this post. The area of distributed data.
We started by running a database on prem. Each branch or store got their own db. 15 years later always-on networking happened. 15 years after that, all businesses have fibre.
So now all the branches use a giant shared online database. With merged data. Uuid based this task would be trivial. Bigint based, yeah, it's not.
Along the same timeline data started escaping from our database. It would go to a phone, travel around a bit, change, get new records, then come home. Think lots of sales folk, in places without reception, doing stuff.
So you're right in the context of a single database (cluster) which encompasses all the data all the time.
But in the context where data lives beyond the database, using uuids solves a lot of problems.
There are other places as well where uuids shine.
So as with most advice when it comes to SQL, I'd add "context matters".
If you're copying a DB, mutating, then merging back in, you just have to reset the bigint pkeys. I can see how in some contexts that might be less convenient (or if merges are very frequent and reads are not, less performant), but that's a special case and not something to assume from the start. For example I've done merges like this before pretty easily with bigints, and I've also been in places where they start out with uuids pkeys then never benefit.
Renumbering bigint primary keys, so as the effect a one-time merge, becomes substantially less trivial if the desire for minimal downtime, coupled with hundreds of related tables, and tens of sites are in play.
In-between is a non-trivial renumbering step, which takes measurable time that invalidates all existing backups.
By contrast uuid based databases do not need this step, and all existing data (some steady distributed, some in backups etc) remain valid.
Depends on your access pattern, you may prefer the other way, even on the same DBMS.
[1] https://github.com/paralleldrive/cuid2?tab=readme-ov-file#no...
In postgres for example, full_page_writes (default on, generally not safe to turn off unless you can be sure your filesystem can guarantee it) means you have to write the entire page to WAL if you write one record. This will make your WAL grow way faster if you're doing random IOs. So right off that bat that's going to be a huge write impact.
In proportional fonts, underscores are generally wider than spaces, creating larger gaps between the underscore-separated parts than between the surrounding space-separated words. E.g. in "AAA BBB_CCC DDD", "AAA"/"BBB" and "CCC"/"DDD" are closer together than "BBB"/"CCC". In some fonts the difference is quite substantial. This makes for incorrect/unintuitive visual grouping.
You have to press Shift to type them. On mobile keyboards, underscore is usually one extra layer removed. For voice dictation, it's also longer than "dash" or "minus".
Jesus, what a nightmare.
The solution to your entity problem should be the same. You do the reasonable, practical thing, and rename/refactor if they drift away from the original mental concept.
Exactly right. You've succinctly stated the biggest problem with almost all modern database design.
> So do you not name your tables?
I do, but that name does not represent a Type of Entity, where all entities therein are Exactly Thus, and all entities everywhere else are Absolutely Not At All Thus. Instead, it represents a statement I want to make about entities. Any "natural" meaning you put into your identifiers about what they are is defeating the point of the identifier.
> You do the reasonable, practical thing, and rename/refactor if they drift away from the original mental concept.
And now all of your IDs that are "in the wild" have expired. Can I still submit a request using the old ID, before you renamed Employee to WorkPerson and then to MobileLivingBeing and then to PossiblyMobilePossiblyLivingBeing? And it's not "if" they drift away, it's "when". And it's not just that they change over time, it's that they change from one perspective to the next. You can never have two distinct disciplines of the business ever referring to the same entity, because they don't agree on what the types mean. That bears repeating: you can never have two different disciplines both referring to the same entity unless they agree on what the types are, and they don't, because their terms have different meanings. Do your accountants and your maintenance people and your capital planning people and your corporate leadership all agree on exactly what a "facility" is? Because if they don't, they literally cannot even refer to the same entity. Good luck with your microservices.
aws_access_key_id = AKIA367COJQOEU3UOE
aws_secret_access_key = a7Ed0F80a0AF6606/MQG3+4o/o
It's frustrating that you can select the access key by double clicking it but not the secret access key because of those / characters.My addition for your consideration: https://github.com/mik3y/django-spicy-id
But all these were talked about and considered before it was punted to a later time. https://github.com/uuid6/uuid6-ietf-draft/issues/27 https://github.com/uuid6/new-uuid-encoding-techniques-ietf-d... https://github.com/uuid6/new-uuid-encoding-techniques-ietf-d... https://github.com/uuid6/new-uuid-encoding-techniques-ietf-d...
But there is always TypeID in the meantime which uses UUIDv7 under the hood: https://github.com/jetify-com/typeid
Either way, I am in favor of prefixing and using alternative encodings, but it will need some time to figure out the best route. In the mean time, there are so many alternatives. TypeID, NanoID, ULID, etc. I even made my own quick one just for giggles: https://github.com/daegalus/snowflakes
Clustered index with random data stored as chars as the PK, what a great time! You will surely not regret this decision later.
The author also compares to Stripe tokens, which is a strange comparison as you can see they also have a time component towards the beginning.
For those who think this doesn’t work in distributed systems, it absolutely does – PlanetScale uses them internally [0]. If what is likely the largest MySQL (under Vitess) cluster in the world can manage, yours can too.
If this is still untenable, then anything k-sortable (like UUIDv7, as the sibling comment mentioned) is a vast improvement over randomness. Don’t cause B+tree page splits, especially in an RDBMS with a clustering index like MySQL.
[0]: https://github.com/planetscale/discussion/discussions/366
It’s still a good idea, IMO, to think about table design starting as though you have a natural key (composite or singular). It helps develop the schema; you can then drop in a serial/identity/autoincrement column, and use the other relationships as FKs.
WHERE id = 123
compared to WHERE site = 'HN' and username = 'hot_grill'
The last query is easier to write when you are querying the database manually, but I find the first more easy to handle programatically. It's easier to pass an argument from an URL or a message queue in this case.As systems evolve, you may find that you need a third component to the natural key. If you don't use a simple id, you need to update every query that references the natural key
Hexadecimal is safe.
Which is why the standard has been base16 (or base10) for so long.
/s… but only halfway
I usually favour b32 for IDs. There's also word encodings, but frankly those have more lewd combinations than they don't to a mind such as mine.
No it's not.
Edit: downvoters, a tame example is 0x72b5473d5a567200.
Honestly, if seeing the character string "b00b" is a problem, you have bigger problems.
Oh yeah?
ABADBABE B16B00B5 0B00B135 BEEFBABE CAFEBABE DEADBEEF
And a few others that I'm probably forgetting...
Nobody does this. Normal users don't even know this is a thing. I worked on an app that did something like this for placeholders in generated text, and in all our extensive testing and high-touch rollouts, we never saw anyone use it.
It's nice that you took the time to think about it, but it's not that important.
Double-clicking is just one aspect of the wider principle, that is you make life easier for heuristics of all kinds.
What was the thing your product did?
Sort of like:
The applicant is fully employed as a _PROFESSION_ at _NAME_OF_COMPANY_
Our designer excited told us how he specifically used underscores so you could double-click the placeholder and just type over it.Quite literally not worth the trouble.
There is even a specific Microsoft SQL-time-ordererd UUID format which is sorted after byte shuffling..
We store ULID in binary(16). Works nicely. Only difference from UUIDv7 is the version bits..
I found today a PHP UUID library that does one they call UUIDv8 which has more time (like ULID)
They are less likely to take a ULID and store it in binary(16). (Might store it in char(26), which is less storage efficient but is still sorted.)
Apart from that I happen to LIKE the fact that I can very quickly see the difference between a random identifier and a sorted identifier, they have quite different characteristics after all.. Although I guess people might eventually get used to the initial bytes being 0 on UUIDv7, or just learn to recognize the version bytes..
On that topic, why are the 8 fixed bits of a UUID not concentrated on the same byte. Perhaps the first one. Huge mistake..
In the case where it's necessary to have a text representation (e.g. in some user interface) I guess it's fine to choose whatever (stable) transformation you like, but the standard ways specified in RFC-4122 (hexadecimal, with or without hyphens) seem like the most foolproof. Regarding the logs search use case, the first "chunk" of a uuid is usually well more than enough for a unique match, IME.
Also, I've been burned before in cases where some clever transformation was used to make a uuid look different in text form, because in order to synthesize the actual binary uuid I first have to reverse engineer the transformation--e.g. to find the database record corresponding to some http request log message. That's just annoying, and the polar opposite of "user friendly" for the user story of an engineer trying to figure out what's wrong with the system.
So.. I guess my vote is to stick to the standard.
Thank you. This is overthought so much, including with partially-random things like uuid3, 5, 7.
Like virtual machines or embedded systems.
It's more like, you used id X in your db and 8 months later another X lands, after billions and billions of rows have been inserted...
The coarsest encoding to have this property is Base32 where it remains easy to memorize first few letters without needing to memorize case.
But like... why? This article literally does not explain the benefits beyond copying, they are just assumed. I'm not immediately sold on shorter === better, especially when the updated UUIDs are only marginally shorter and you have now introduced the overhead of a translation layer for one of the most basic building blocks in your application.
Let's say you have a customer with that UUID as their ID. Do you expect them to recite their UUID perfectly to you every time? What if they made 5 transactions, each with their own UUIDs and you need to look them up, do you now expect them to read out 6 fairly unwieldy IDs?
The article is about the UX of UUIDs. Yes, there's a translation layer and more dev work to implement, but the shorter size and use of non-ambiguous characters is a massive improvement in the usability for the end users.
Also, you don't explain why copying isn't good enough, you just reject the article's reasons
Given:
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua.
I'd like to, say, double click on consectetur to select it (which works), bt then, while holding Shift, I would like to double click on elit so that the selection of consectetur is preserved, and extended over elit, including adipiscing.
The behavior I typically see is that, while holding Shift, the first click will extend the selection to the exact character of elit that I'm pointing at. The second click will then cancel the selection and select all of elit.
Like it doesn't mean a damn thing that I'm holding down Shift!
Ironically, I can make multiple selections that way using Ctrl in Firefox; Ctrl does modify the semantics of the second click while there is a selection.
for display, you have various ways to encode that number into something easier for humans, what I prefer:
- short word - no offensive words - with a checksum (so we easily spot any copy paste mistake) - not sequential - can be put in a url without extra encoding
this is our implementation of that (base32 and luhn code)
Triple click can be used.
https://softwareengineering.stackexchange.com/questions/3855...
Maybe the browser (and other) UI is wrong about word boundaries.
Use immutable human-readable identifiers like "slugs" and/or "natural keys" in addition to robust primary keys.
> TLDR Please don't do this: https://company.com/resource/c6b10dd3-1dcf-416c-8ed8-ae56180...
That URL is fine, it should just be a 30X to https://company.com/stuff/cute-slug and you should use the user-friendly URL where possible.
There is no "central coordination" required to set up URL redirects, by the way.
The PGP Word List for translating hex into words: exists.
> They provide a reliable way to ensure that each item, user, or piece of data has a unique identity.
It is the registration of a UUID in a database which prohibits reuse that does that. If you aren't do doing that, ensuring that each use of a UUID is not a reuse of an already assigned UUID, they are not UNIQUE.
The whole point of using UUIDs is that you can generate them locally without central coordination--if you want to coordinate your identifiers, you can use a much friendlier ID length (which is explained in the article).
So any UUID coming from an untrusted source (like a client application) should be checked for uniqueness. However your client apps can be written assuming that their randomly generated UUIDs never have collisions.
The alternative of having some sort of "placeholder ID" until the sever gets back to you (or you get back online) adds a lot of client complexity.
Of course you have to deal with the implications of people intentionally colliding UUIDs, so maybe don't generate them client-side.
a reason to use uuid (eg v4) is that you can generate id's distributed without fearing collisions. it can happen, but is not likely.
so the uniqueness comes from a property of the ID and not the database.
"Reducing the length of your IDs can be nice, but you need to be careful and ensure your system is protected against ID collissions. Fortunately, this is pretty easy to do in your database layer. In our MySQL database we use IDs mostly as primary key and the database protects us from collisions. In case an ID exists already, we just generate a new one and try again. If our collision rate would go up significantly, we could simply increase the length of all future IDs and we’d be fine."
No, even that is not true. If all digital (and non-digital) storage media ever manufactured by humans - meaning all hard drives, tape drives, CDs, DVDs, BluRays etc. ever manufactured to date, and every book and word ever printed or written down. If those ALL were only filled in with UUIDv4s generated from a good random source .. you would still not see even one single collision!
UUID collisions are only possible with currently known human technology if your randomness source is not good enough. And it will remain so unless there are some astronomical leaps in digital information storage technology - at least 10 orders of magnitude more storage than currently exists.
EDIT: I thought of a way for programmers to mentally visualize how unlikely UUID collisions really are. Let's imagine that in some not-too-distant future, there are 10 billion people on Earth. Each of them are given one thousand CPUs. These CPUs have 1024 cores each, and they run at 10 GHz (clock cycles per second). The CPUs implement a hypothetical instruction that can generate a totally random UUID in one clock cycle.
As an experiment, all people on Earth one day decide to program all their thousand CPUs each to run a tight loop that will indefinitely generate UUIDs on all 1024 cores and then immediately discard them.
After continuing to run this experiment (whose electicity bill will make Bitcoin look like Earth Hour) all day, 24/7, for about 800 years, the likelihood of one UUID ever having been generated twice will have exceeded 50%.
- 10 billion people =~ 2^33
- 1000 CPUs =~ 2^10
- 1024 cores =~ 2^10
- 10 GHz =~ 2^33
So: one second's computation by all of these people is 2^86 UUIDs generated. UUIDs are 128 bits. With probability essentially 1, there will be a collision within one second.
The reason is known as the birthday paradox. If you sample random values from a set of size k, after you've chosen about sqrt(k) values you will have chosen the same value twice with probability very close to 1/2. By 10*sqrt(k) samples you'll have found a collision with probability well over 90%.
In this case, after sampling 2^64 values you'll have a collision with probability 1/2. That happens in roughly 250 nanoseconds (2^-22 seconds) in your thought experiment.
2^64 sounds like a lot, but in many contexts it's not all that much. Every bitcoin block mined takes well in excess of 2^70 SHA evaluations. Obviously the miners are not dedicated to generating UUID collisions, but if they were they'd easily find thousands of them in the time it takes to mine one block (this neglects the fact that it is much easier to sample a UUID than to evaluate double-SHA256).
Right. Forgot about that little thing. You're absolutely correct.
With this taken into account a _single_ person with 1000 pieces of 1000-core, 10 GHz CPUs generating UUIDs will generate a match in a few minutes.