You Don't Need UUID
henvic.dev
henvic.dev
With sequential IDs, the database state is pretty rigid and cannot easily be restored while still in continuous use, because you can relatively easily run into primary key collisions/violations that can make it incredibly difficult to backup and restore data without taking the whole database offline for writes. Things like scaling out to multi-region and sharding also can get pretty quickly problematic - you end up with duplicate IDs that makes it impossible to easily move data between regions - unless you use a centralized unique ID service (see Twitter's Snowflake for an example). UUIDs just avoids this issue entirely and you can move data around with ease.
With UUIDs, things like being able to grab some test data from production, scrub it, insert it into your local database, test with it, and then update it back to staging or production uses trivial normal insert/update operations.
With sequential IDs, this task becomes a nightmare of ID collisions and manually altering foreign keys if you do run into duplicate IDs.
If anything really does need to be exposed to end users, I typically use another unique string field like username, post slug, etc. rather than directly exposing the UUID to search engine traffic, while still using UUIDs internally everywhere within my database.
The author is explaining the drawbacks of UUID as a user experience problem.
UUIDs are built for machines.
The solution to providing an identifier for humans doesn't need to involve long—difficult to remember—numbers at all.
There are many better alternatives to UUID for the general case in which UUIDv4 is used, that provide some of the following advantages (based on the implementation): - Simpler spec and implementation - Easier to enforce compatibility across implementations (no ambiguity about letter case, order or hyphenation) - Smaller size and better readability - Lexicographically sortable based on time of generation - Deterministic uniqueness guarantee
UUID is too much of a jack-of-all-trades-master-of-none. It originally defined many versions (which should have been called "formats") for flexibility, but nowadays only UUIDv4 (and less often UUIDv1) seem to be in use. UUIDv3 and UUIDv5 rely on the insecure MD5 and SHA-1, which, while it seems UUIDv2 and Nil UUIDs were rarely ever used.
I hate it when developers are always reaching immediately for UUIDv4 while they almost always have a better choice for their case.
There is also no reason to give up 6 reserved bits of a uuid v4's 128 bits (it's only 122 random bits + 6 bits of unnecessary version info). If you want random IDs, make your own. Simply generate 16 bytes of raw data; or combine two random 64-bit integers; or combine a timestamp prefix with random bytes at the end. Your needs probably don't warrant exactly 128 (well, 122 bits) that uuid v4 gives anyway, so you can customize to a specific number of bits. You can also use base62 (0-9, A-Z, a-z) or base64 instead of hexadecimal (0-9, A-F) to reduce the number of displayed characters (eg. in URLs), while omitting the stupid hyphens of a uuid too.
tldr; uuid v4 shouldn't even exist, and certainly should not be used unless integrating with pre-existing systems.
Could you name a few with their trade off or point me to some resources? I only know about UUIDs and auto-incrementing sequentail ID, but I would love to add a few tools to my toolbox.
PostgreSQL (and other RDBMS) also supports storing and indexing the data in UUIDs in a binary format. With binary, the space and speed concerns are greatly minimized, especially when comparing UUIDs to other randomly generated homegrown strings.
No doubt this phenomenon successfully plucks at some emotional vulnerability most of us have, and thus brings in the clicks.
edit: I also disagree with most of the points this article makes, but that’s just, like, my opinion, man.
UUIDs offer by far the best developer UX. They're first class objects for most of your tech stack, and they eliminate or massively simplify a lot of potential backend issues, from idempotency to merges to sharding.
The main downsides are two:
- They're not human friendly. This is a good thing. Surrogate keys should never be exposed to the human user, because what is a 1:1 mapping today may not be 1:1 any longer tomorrow. If you need to let the user input a unique object reference manually, you're free to design that reference with human UX in mind. Maybe it's a set of dictionary words, maybe it's a code with an expiration date, maybe it's the same code but referring to different versions of the object depending on the user.
The only people who may need to copy/paste UUIDs manually are database admins, and they will thank you so much more for never bothering them with fixing a pkey clash or an accidental join on the wrong key.
- They are less performant than sequential integers. For 99% of projects this is not going to be an actual issue. Once you _do_ find that your queries are being slowed down by the size of primary keys specifically, you have my absolute blessing to add an integer sequence column and make that the primary key.
I sometimes have this odd cognitive dissonance about it and have to consciously remind myself that this intuition is ridiculous.
Is it just me?
A superficial post. Doesn't talk at all about when/when-not to use source-generated random identifiers.
The only argument I can think of in favor of this proposed alternative is if at some point in your business workflow someone needs to manually type an ID and cannot copy paste it. Sounds pretty rare to me. Even support is now done more and more via chat systems.
Two cases:
- screenshots - a page that's on your phone that you want to see on your computer
Sure, those are not the most frequent, but I'm glad I don't have to type a UUID each time.
The thing I don't like about UUIDs is that they're not necessarily unique across multiple universes.