Analyzing New Unique Identifier Formats (UUIDv6, UUIDv7, and UUIDv8) (2022)
blog.devgenius.io
blog.devgenius.io
* Construct a tuple of the current TAI64NA time, the process ID, some other IDs, and a sequence number.
* Encrypt this 384-bit tuple with a 1024-bit secret using SURF.
* The 256-bit result is the ID.
https://github.com/bruceg/ezmlm-idx/blob/master/lib/surf.3
One could adapt this approach to UUIDs.
However, for object retrieval (e.g. "show me my accounts with object IDs X, Y and Z"), you shouldn't be just relying on the unguessability of the ID anyway - you should be using appropriate access control permission checks (because for object IDs even if they are unguessable they are frequently not guaranteed to be private, e.g. they show up in log files, they may be in URLs that people share, etc., and they don't expire, unlike a session ID).
A frequent source of severe security vulnerabilities are bugs in these access control checks, which can lead to insecure direct object reference (IDOR) vulnerabilities, where I can say "Show me account with ID X" even though I don't own or have permissions to that account. If you're using incrementing integer IDs for your object IDs here, abusing this bug is trivially easy - just walk up and down the IDs next to yours. I think it was the Parler app that had a particularly egregious case of this - they weren't doing any access control checks on posts at all.
However, if your object ID is a v7 UUID, this bug becomes much more difficult to abuse, because your IDs may not be completely unguessable to a strong adversary, but you should be able to detect the failed access attempts when you see someone trying billions of access attempts and only getting not found errors (of course, you should be rate limiting these attempts in the first place).
The only issue is lack of trusted implementations. But it’s easy enough to implement it.
This makes ULIDs easily predictable and defeats one of major reasons to use UUID. 1 millisecond is a lot and thousands records could be generated in one millisecond, especially with some kind of batching.
This algorithm could be improved: either generate new random 80 bits until they're higher than previously issued ULID or select few bits for counter and rest of bits for randomness. But then we'll end up almost with UUIDv7, I guess.
Also their representation idea... Well, I think that UUID representation is a widely used everywhere.
So basically I think that ULID is an amalgamation of questionable ideas and I'd stay away from it.
Note that the UUID v7 spec is largely modeled after the ULID spec. ULIDs came first, and they traded the "standard" UUID format of 8x-4x-4x-4x-12x hex string for the more compact Crockford base32 format, and there are some other minor differences in number of timestamp bits vs. number of random bits, but they are otherwise functionally equivalent.
Picking something primarily because of its human-readable form rather than its machine-readable form is probably not the best decision criterion.
ULID on the other hand depends on an uncommon encoding scheme. It is more complicated than hex encoding, so it's easy to make mistakes when implementing it.
Finally, the encoding scheme is weird, because it "wastes" a few bits at the end, so if you are not careful enough you could end up in a situation where two ULIDs in text form are different but once you convert them to bits they are the same.
UUIDs rarely have to be unique compared to every other place in the world where they are used, they just need to be unique within the same context. The reason for using NIC addresses and time is to ensure that the chance of a collision is really small (but not 0). Saying that you are making them a little more unique doesn't sound like good motivation.
The second part is that just using random numbers gives you no real guarantees of uniqueness since without seeding with time and NIC, it would be highly likely that multiple systems around the world would generate the same number.
"Highly likely"? No. You are ignoring probability. See UUIDv4, which does not use time or NIC/MAC but does use 122 random bits.
https://zelark.github.io/nano-id-cc/
At the default 64 character alphabet with a 21 character key length it would take ~41 million years in order to have a 1% probability of at least one collision if you generated 1000 ids per second.
Some info from that page:
- Safe. It uses hardware random generator. Can be used in clusters.
- Short IDs. It uses a larger alphabet than UUID (A-Za-z0-9_-). So ID size was reduced from 36 to 21 symbols.
- Portable. Nano ID was ported to 20 programming languages.
Proof of publication in a fully decentralized academic journal, as an alternative to the current grifting scientific publishing industry.
As a consensus protocol for a distributed graph database, where you ensure consistency through exclusive ownership of nodes by single writers.
But those things don't offer any get rich quick shemes, so we're stuck if crypto bros and pyramid shemes.