TIL: Versions of UUID and when to use them
ntietz.com
ntietz.com
Only no known details if the only document you're reading is the notoriously poorly-specified RFC. Here you go: https://pubs.opengroup.org/onlinepubs/9696989899/chap5.htm#t...
There are also “version 0” UUIDs that you are very unlikely to ever come across but should be noted because they are the source of the reserved bits (via wastefully setting aside an entire octet for Address Family) that later allowed the other “versions” to be specified in a compatible way. Read my research about them here in my UUID library: https://github.com/okeeblow/DistorteD/blob/NEW%E2%80%85SENSA...
I decided to support them Because It's Cool™ but still need to figure out how to handle the date rollover of them and the even-older Apollo UIDs:
irb> ::GlobeGlitter::from_ncs_time
=> "#<GlobeGlitter 40639cd25341.02.00.00.e0.4c.18.00.69>"
irb> ::GlobeGlitter::from_ncs_time.to_time
=> 1988-12-21 14:52:02 UTC
irb> ::GlobeGlitter::from_aegis_time
=> "#<GlobeGlitter 00000000-0000-0000-4814-17c8b0080069>"
(Proper AEGIS `#to_str` not implemented yet lol)To be fair, in RFC9562 I did cite two documents that are UUID Version 2 specifications. But RFC4122 was too cryptic for my taste.
As for the historical UUID types specified by the 0-7 Variant space: We are starting work on an informational RFC that will help folks understand those.
See https://github.com/yocto/draft-yocto-uuid if you want to add to some discussions and/or review text. We are still a bit early into the stages but I hope to have some progress soon.
I found the details in about 2 minutes: Click the link in the article to take me to the section of RFC 9562 that says it's defined as part of DCE, click the first link in that paragraph to go to the spec, ctrl-f "UUID", then jump to appendix A (deceptively named "Universal Unique Identifier") which has all the details.
Is it really too much to ask to CLICK YOUR OWN LINKS?
i enjoyed reading the appendix though as a snapshot of time.
I agree the sentence is a bit unclear, but I don’t think it’s misleading or whatever.
and then a small hash with base64 or 37 or whatever is in vogue these days.
thats what old timers used before uuid 1.
guess we should guerilla standardize something like this as uuid-0 or uuid-deprecated-2.0 for keeping up with the spirit.
The problem with auto-increment integer ids is they are not always possible with distributed systems.
I'm fairly certain the first rule of websec is you never trust the client. I definitely would not trust a user's browser to directly insert a value into a DB.
> What if you don't know how many nodes you have?
Shouldn't matter; you have a centralized system that hands out chunks of IDs on-demand (and has its own mechanism to ensure no repeats). This is similar to what Vitess [0] does.
[0]: https://vitess.io/docs/20.0/reference/features/vitess-sequen...
Not every piece of information is confidential in every system. Sometimes a UUID is just that, a UUID.
> you have a centralized system that hands out chunks of IDs on-demand
I don't follow. If your system requires a central node that can reliably generate unique auto-incrementing integer IDs, why bother with UUIDs at all? Just base-64 encode the integer ID, or hash it with a salt to protect against enumeration attacks, if you want.
If you don't want the dependency to a centralised system, just use UUIDv7, which is just a timestamp plus random bits, or implement a shorter version of it. There is no need to overengineer.
I also don’t follow. I thought your initial assertion was that auto-incrementing integer IDs weren’t always possible, thus the need for UUIDs.
Monotonic ints, or more broadly anything k-sortable, are generally optimal for RDBMS indices due to most indices being B+trees. That’s why there’s such enormous effort towards NOT using UUIDv4.
> just use UUIDv7
Indeed; this is my recommendation when devs insist they can’t possibly use integers. Personally, I maintain that most places can use ints, it’s just that they’ve hideously over-complicated things to the point that it would be far too much work.
You suggested timestamp+autoinc, and my initial assertion was that auto-incrementing integer IDs weren’t always possible, thus the need for the random part after the timestamp (a la UUIDv7). I see that we have actually been on the same page.
The real benefit of UUIDs is the 'consistency' of the one-size-fits-most approach. If you can do without IDs humans can read out, or readable plain text logs, or compressibility, or recognisable formats for different types of ID? Then UUIDs can be used for anything from customer orders to web requests to log lines.
I wouldn't use it for assigning bank account numbers, but for most web or app stuff it's fine.
I think the reason there's no other popular standard is you give up something. 128 bit gives a pretty low risk of collisions in almost all uses, but as you go smaller you start having to consider the specific scenario and impact, etc, which doesn't work well for a standard.
You could use another encoding (eg base64 or base85) to get it shorter, but you start sacrificing other things (case sensitivity, url-safeness) - again, not great for a standard.
UUIDv7 will only work until year 4147 (compared with ULID's 10889AD), but by then I think we'll have another UUID version we can switch to.
Here's my implementation in Python: https://codeberg.org/prettyid/python, https://pypi.org/project/prettyid
And a rudimentary TypeScript library: https://codeberg.org/prettyid/js, https://npm.im/prettyid
It's the same UUID just in 22 character form and can be converted back. It's n ot really a conversion because a UUID is just a 128 bit value so its an alternative representation.
483971cf-aad7-4c84-abf1-4a94c9d72f99 -> SDlxz6rXTISr8UqUydcvmQ (length: 22)
fb67926f-3cfb-486c-a7da-30662147a20b -> A2eSbzz7SGyn2jBmIUeiCw (length: 22)
799069a9-b32a-415f-b689-a8cc3f51bfa4 -> eZBpqbMqQVA2iajMP1GBpA (length: 22)
8161ee0b-f7a5-4b32-95ea-9b9efe94e5f2 -> gWHuCBelSzKV6pueBpTl8g (length: 22)
b1ea416c-f209-43cb-bfaf-d9cf6229459e -> sepBbPIJQ8uBr9nPYilFng (length: 22)
ee70989a-b614-4665-9881-41054544c313 -> 7nCYmrYURmWYgUEFRUTDEw (length: 22)
cce06fe2-b64f-47bc-a91a-d3dfd343e1e5 -> zOBv4rZPR7ypGtPf00Ph5Q (length: 22)
aea3de6e-e769-4c8d-ba2d-77922d227176 -> rqPebudpTI26LXeSLSJxdg (length: 22)
import uuid
import base64
def make_short_uuid(data):
encoded = base64.urlsafe_b64encode(data).rstrip(b'=').decode('utf-8')
return encoded.replace('-', 'A').replace('_', 'B')
def generate_and_print_uuids():
for _ in range(8):
uuid_obj = uuid.uuid4()
uuid_bytes = uuid_obj.bytes
print(f'{uuid_obj} -> {make_short_uuid(uuid_bytes)} (length: {len(make_short_uuid(uuid_bytes))})')
generate_and_print_uuids()If you really want to not have _ or - in your short form UUIDs you could just discard the UUID when you create it if the short form includes those characters and try again.
c22c1dcf-ea74-470e-acbf-b1722e243025 -> wiwdz-p0Rw6sv7FyLiQwJQ (length: 22)
Reversed: c22c1dcf-ea74-470e-acbf-b1722e243025
8702aecb-6d09-4a5e-8cc8-621aada6ed96 -> hwKuy20JSl6MyGIarabtlg (length: 22)
Reversed: 8702aecb-6d09-4a5e-8cc8-621aada6ed96
643a9829-9f91-4b88-80a6-db2c0eb83e8b -> ZDqYKZ-RS4iAptssDrg-iw (length: 22)
Reversed: 643a9829-9f91-4b88-80a6-db2c0eb83e8b
8e7f3c1a-3d19-425e-8803-fe55a296688e -> jn88Gj0ZQl6IA_5VopZojg (length: 22)
Reversed: 8e7f3c1a-3d19-425e-8803-fe55a296688e
1859f017-a5f3-4875-825a-fdd531dfac1a -> GFnwF6XzSHWCWv3VMd-sGg (length: 22)
Reversed: 1859f017-a5f3-4875-825a-fdd531dfac1a
6a153b44-7fca-45b2-b13e-7f45790be7bf -> ahU7RH_KRbKxPn9FeQvnvw (length: 22)
Reversed: 6a153b44-7fca-45b2-b13e-7f45790be7bf
fd6bad83-a0f8-4c7f-baf1-10374be3e8e9 -> _Wutg6D4TH-68RA3S-Po6Q (length: 22)
Reversed: fd6bad83-a0f8-4c7f-baf1-10374be3e8e9
cf2452d4-947b-4b92-a280-ff869e77ba65 -> zyRS1JR7S5KigP-Gnne6ZQ (length: 22)
Reversed: cf2452d4-947b-4b92-a280-ff869e77ba65
import uuid
import base64
def make_short_uuid(data):
return base64.urlsafe_b64encode(data).rstrip(b'=').decode('utf-8')
def reverse_short_uuid(short_uuid):
# Add padding back to make it Base64 decodable
restored = short_uuid + '=' * (-len(short_uuid) % 4)
# Decode the Base64 string back to bytes
return base64.urlsafe_b64decode(restored)
def generate_and_print_uuids():
for _ in range(8):
uuid_obj = uuid.uuid4()
uuid_bytes = uuid_obj.bytes
short_uuid = make_short_uuid(uuid_bytes)
reversed_uuid_bytes = reverse_short_uuid(short_uuid)
print(f'{uuid_obj} -> {short_uuid} (length: {len(short_uuid)})')
print(f'Reversed: {uuid.UUID(bytes=reversed_uuid_bytes)}\n')
generate_and_print_uuids()IMO we need to be clear on the distinction between (A) the UUID bit-generation scheme versus (B) the way it is encoded for human use/reading/transcription.
They are mostly-separate problems.
For example, you could have a very secure mathematical scheme, but it gets ruined by a horrible representation where each bit is written as either a capital-I, a lowercase-l, or the number 1.
Conversely, could have a deeply insecure scheme that uses a nice compact serialization where everything is grouped into chunks and "1Il" confusion is not possible and there's a check-digit, etc.
You can the discussions here: https://github.com/uuid6/new-uuid-encoding-techniques-ietf-d...
By reading the Wikipedia page I'm failing at understanding why we invented something called universally unique identifier and have different types of it, some of which can be traced back to the original pc. Is it because mixing some Mac codes increase the chance of the uuid2 being randomic or does it have a different reason? For privacy reason, could we just not have a very long identifier with many different chars to choose from so that we have so many combinations that we're almost guaranteed we're using non duplicated uuids?
I think a lot of the confusion can be traced to the very earliest AEGIS implementation where the Apollo engineers started using “canned” (their term, i.e. static or well-known) UIDs to identify filesystems. Over time the popular usage of UUID fully shifted from ephemeral identifiers where duplicates were intentional toward canned identifiers where duplicates were unwanted and the two dimensions were random-and-also-random.
The history gets even more complicated because Microsoft hired one of the top Apollo guys to do MSRPC for Windows NT, so there is also “GUID” which differs from UUID in the layout of the fields and is not mixed-endian despite what a lot of sources will tell you. In addition to ephemeral RPC message-identifying GUIDs Microsoft are also in love with canned GUIDs for identifying COM classes, media codecs, and almost anything else that would ever need a well-known identifier. See https://gix.github.io/media-types/ for example.
Apologies for linking my own repo twice in the same comment section but I started (and need to get back to) compiling the history of all this in the README of my UUID library. Apollo started in 1980 and the Leach/Salz UUID RFC draft didn't happen until 1998 so there is a huge amount unsaid by the modern standards: https://github.com/okeeblow/DistorteD/blob/NEW%E2%80%85SENSA...
Like in Go, it's just uuid.New().String() vs using crypto/rand to read random data, convert it into Base64 of hex... which will take more lines and effort.
I propose leftpad.js
Even if you do just want a random identifier (not really the original point of UUID but has become their most popular form) I still think it's cool how random UUIDs have a little flag bit to tell you that it's intended to be random. Useful when one runs across a lone identifier with zero context.
e.g.
XXXXXXXX-XXXX-4XXX-[89AB]XXX-XXXXXXXX
From looking at all the ones in my system.
Whether that's actually worth anything for a particular use case is a good question, and the answer will mostly be "not just no but HELL NO!"
Though I am working on a way to solve this problem with UUIDs beyond 128 bits so we don't have to truncate the hash.
Otherwise, you'd want longer outputs.
Cue the security experts who say otherwise…
The python uuid standard library doesn't have V7 yet, and there is a package called uuid7 which is unmaintained, and not in compliance with the latest standard. That's using nanosecond time precision rather than millisecond, which means the leading bits are larger than they are meant to be.
If you use that unmaintained uuid7 package and later change to the correct implementation your uuid7 will go backwards, which is a breaking change considering that monotonicity is a key property of uuidv7.
Also while technically true - it could technically break monotonicity (records added in the same nanosecond could be out of order) they'll still be all "near the end of the file, likely in the same page" such that performance implications are negligible.
As a general rule I would avoid any program making any assumptions about a uuid. Programs should treat it as an opaque binary random value. Doing so avoids any future incompatibilities.
(You can use ULID's presentational tools with UUIDv7, though.)
Yeah, I've gotten in the habit of stripping hyphens from the string representation of UUIDs in a lot of the code I write for that reason.
https://developer.mozilla.org/en-US/docs/Web/CSS/user-select
You can control this behavior in CSS with `user-select`. Peep my fiddle: https://jsfiddle.net/gLyph5km/
(ULID representations also are shorter because they use a wider character set, which is nice though not critical.)
I did that, works pretty good: https://codeberg.org/prettyid/python
Some more context in a sibling thread: https://news.ycombinator.com/item?id=41355218
Perhaps you mean something like "standardized hash of all columnar data for the table row," but then you're just reinventing elasticsearch/lucene, with all its pros and cons. The power of foreign keys for a RDBS is that they are pointers, and as pointers, the mutability of their underlying data is what makes them powerful. I think I get what you're asking for, but I also think there can be no possible standard that is reasonable unless you have the technology to take a total snapshot of the universe, at which point, why not just measure the universe itself as your database? Perfect storage system.
Their is a reason cybersecurity or UI/UX or product design isn't always left to the developer. The coder write code that fits certain criteria they are given, then someone down the line might QA check it, fuzz inputs or security review the code. How well this is done depends on the product,market, and environment.
In most cases creation time is not sensitive. Therefore, for most cases uuid7 is the best trade off currently.
i hate how long it is
something like youtube URLs but guaranteed to be without duplicates
Maybe one way is to split up a random assignment space and assign to each distributed node, but that would be more complex.
You might think your "Uber, but for short term giraffe rental" startup needs that sort of guarantee in the investor demo prototype, but it doesn't. Just use an auto increment in Postgres or MySQL (or an integer primary key column in SQLite). If you fool those investors into pouring the money pipe over you, your first real technical/senior engineer hire is gonna throw all the code you have running on your laptop away anyway.
Maybe someone at Meta sat down once and figured "we have 4 billion users, each of who on average has 3 cats that they take a dozen picture of every day, so about 150 billion cat pictures per day. So if we name them using uuids we're still good for almost 50 thousand years before we have a 50% chance of displaying a pic of Fluffy when we should have displayed a pic of Mr Whiskers". Then they promptly ignored the problem (or fixed Zack's code that was using php's hash("md2", $query['catname']) ).
But yes, if you're just using them as primary keys in a database you're probably fine with auto increment for most use cases.
Then just encode it in a format of your choosing. Youtube uses a modified base64 encoding (no padding, and + and / are replaced by - and _). And youtube video ids seem to also be 64 bits, just like xtea output.
If you're distributed, look into vector clocks[1] or snowflake[2][3]
[1] https://en.wikipedia.org/wiki/Vector_clock
[2] https://github.com/twitter-archive/snowflake/tree/snowflake-...
As long as they gave sufficient randomness etc, from a program perspective they are unique id's.
There are already multiple versions in active use (4, 7 and arguably 8) so you really shouldn't be using the uuid as anything but a long-random-value.
Yes, the database engine may appreciate one version over another for performance reasons, but that's irrelevant to most developers and programs.
Want visually recognizable unique identifiers?
JAGRSW-UID-<192bit-input-from-urandom-encoded-in-base64>
Need to shave off some bytes? JUID-<192bit-input-from-urandom-encoded-in-base64>
Same byte size as UUIDs, arguably more "secure." Can I become an ACM Fellow for solving this problem now?Seriously, these UUID debates are about as sensible as arguing over XML.
What was important during the times when we didn't know how to generate random numbers on computers, perhaps shouldn't be as important today?
Not getting pwned by incrementing href attacks is good, past that, web scale bro.