I think the reason there's no other popular standard is you give up something. 128 bit gives a pretty low risk of collisions in almost all uses, but as you go smaller you start having to consider the specific scenario and impact, etc, which doesn't work well for a standard.
You could use another encoding (eg base64 or base85) to get it shorter, but you start sacrificing other things (case sensitivity, url-safeness) - again, not great for a standard.
UUIDv7 will only work until year 4147 (compared with ULID's 10889AD), but by then I think we'll have another UUID version we can switch to.
Here's my implementation in Python: https://codeberg.org/prettyid/python, https://pypi.org/project/prettyid
And a rudimentary TypeScript library: https://codeberg.org/prettyid/js, https://npm.im/prettyid
It's the same UUID just in 22 character form and can be converted back. It's n ot really a conversion because a UUID is just a 128 bit value so its an alternative representation.
483971cf-aad7-4c84-abf1-4a94c9d72f99 -> SDlxz6rXTISr8UqUydcvmQ (length: 22)
fb67926f-3cfb-486c-a7da-30662147a20b -> A2eSbzz7SGyn2jBmIUeiCw (length: 22)
799069a9-b32a-415f-b689-a8cc3f51bfa4 -> eZBpqbMqQVA2iajMP1GBpA (length: 22)
8161ee0b-f7a5-4b32-95ea-9b9efe94e5f2 -> gWHuCBelSzKV6pueBpTl8g (length: 22)
b1ea416c-f209-43cb-bfaf-d9cf6229459e -> sepBbPIJQ8uBr9nPYilFng (length: 22)
ee70989a-b614-4665-9881-41054544c313 -> 7nCYmrYURmWYgUEFRUTDEw (length: 22)
cce06fe2-b64f-47bc-a91a-d3dfd343e1e5 -> zOBv4rZPR7ypGtPf00Ph5Q (length: 22)
aea3de6e-e769-4c8d-ba2d-77922d227176 -> rqPebudpTI26LXeSLSJxdg (length: 22)
import uuid
import base64
def make_short_uuid(data):
encoded = base64.urlsafe_b64encode(data).rstrip(b'=').decode('utf-8')
return encoded.replace('-', 'A').replace('_', 'B')
def generate_and_print_uuids():
for _ in range(8):
uuid_obj = uuid.uuid4()
uuid_bytes = uuid_obj.bytes
print(f'{uuid_obj} -> {make_short_uuid(uuid_bytes)} (length: {len(make_short_uuid(uuid_bytes))})')
generate_and_print_uuids()If you really want to not have _ or - in your short form UUIDs you could just discard the UUID when you create it if the short form includes those characters and try again.
c22c1dcf-ea74-470e-acbf-b1722e243025 -> wiwdz-p0Rw6sv7FyLiQwJQ (length: 22)
Reversed: c22c1dcf-ea74-470e-acbf-b1722e243025
8702aecb-6d09-4a5e-8cc8-621aada6ed96 -> hwKuy20JSl6MyGIarabtlg (length: 22)
Reversed: 8702aecb-6d09-4a5e-8cc8-621aada6ed96
643a9829-9f91-4b88-80a6-db2c0eb83e8b -> ZDqYKZ-RS4iAptssDrg-iw (length: 22)
Reversed: 643a9829-9f91-4b88-80a6-db2c0eb83e8b
8e7f3c1a-3d19-425e-8803-fe55a296688e -> jn88Gj0ZQl6IA_5VopZojg (length: 22)
Reversed: 8e7f3c1a-3d19-425e-8803-fe55a296688e
1859f017-a5f3-4875-825a-fdd531dfac1a -> GFnwF6XzSHWCWv3VMd-sGg (length: 22)
Reversed: 1859f017-a5f3-4875-825a-fdd531dfac1a
6a153b44-7fca-45b2-b13e-7f45790be7bf -> ahU7RH_KRbKxPn9FeQvnvw (length: 22)
Reversed: 6a153b44-7fca-45b2-b13e-7f45790be7bf
fd6bad83-a0f8-4c7f-baf1-10374be3e8e9 -> _Wutg6D4TH-68RA3S-Po6Q (length: 22)
Reversed: fd6bad83-a0f8-4c7f-baf1-10374be3e8e9
cf2452d4-947b-4b92-a280-ff869e77ba65 -> zyRS1JR7S5KigP-Gnne6ZQ (length: 22)
Reversed: cf2452d4-947b-4b92-a280-ff869e77ba65
import uuid
import base64
def make_short_uuid(data):
return base64.urlsafe_b64encode(data).rstrip(b'=').decode('utf-8')
def reverse_short_uuid(short_uuid):
# Add padding back to make it Base64 decodable
restored = short_uuid + '=' * (-len(short_uuid) % 4)
# Decode the Base64 string back to bytes
return base64.urlsafe_b64decode(restored)
def generate_and_print_uuids():
for _ in range(8):
uuid_obj = uuid.uuid4()
uuid_bytes = uuid_obj.bytes
short_uuid = make_short_uuid(uuid_bytes)
reversed_uuid_bytes = reverse_short_uuid(short_uuid)
print(f'{uuid_obj} -> {short_uuid} (length: {len(short_uuid)})')
print(f'Reversed: {uuid.UUID(bytes=reversed_uuid_bytes)}\n')
generate_and_print_uuids()You can the discussions here: https://github.com/uuid6/new-uuid-encoding-techniques-ietf-d...
IMO we need to be clear on the distinction between (A) the UUID bit-generation scheme versus (B) the way it is encoded for human use/reading/transcription.
They are mostly-separate problems.
For example, you could have a very secure mathematical scheme, but it gets ruined by a horrible representation where each bit is written as either a capital-I, a lowercase-l, or the number 1.
Conversely, could have a deeply insecure scheme that uses a nice compact serialization where everything is grouped into chunks and "1Il" confusion is not possible and there's a check-digit, etc.
and then a small hash with base64 or 37 or whatever is in vogue these days.
thats what old timers used before uuid 1.
guess we should guerilla standardize something like this as uuid-0 or uuid-deprecated-2.0 for keeping up with the spirit.
The problem with auto-increment integer ids is they are not always possible with distributed systems.
I'm fairly certain the first rule of websec is you never trust the client. I definitely would not trust a user's browser to directly insert a value into a DB.
> What if you don't know how many nodes you have?
Shouldn't matter; you have a centralized system that hands out chunks of IDs on-demand (and has its own mechanism to ensure no repeats). This is similar to what Vitess [0] does.
[0]: https://vitess.io/docs/20.0/reference/features/vitess-sequen...
Not every piece of information is confidential in every system. Sometimes a UUID is just that, a UUID.
> you have a centralized system that hands out chunks of IDs on-demand
I don't follow. If your system requires a central node that can reliably generate unique auto-incrementing integer IDs, why bother with UUIDs at all? Just base-64 encode the integer ID, or hash it with a salt to protect against enumeration attacks, if you want.
If you don't want the dependency to a centralised system, just use UUIDv7, which is just a timestamp plus random bits, or implement a shorter version of it. There is no need to overengineer.
I also don’t follow. I thought your initial assertion was that auto-incrementing integer IDs weren’t always possible, thus the need for UUIDs.
Monotonic ints, or more broadly anything k-sortable, are generally optimal for RDBMS indices due to most indices being B+trees. That’s why there’s such enormous effort towards NOT using UUIDv4.
> just use UUIDv7
Indeed; this is my recommendation when devs insist they can’t possibly use integers. Personally, I maintain that most places can use ints, it’s just that they’ve hideously over-complicated things to the point that it would be far too much work.
You suggested timestamp+autoinc, and my initial assertion was that auto-incrementing integer IDs weren’t always possible, thus the need for the random part after the timestamp (a la UUIDv7). I see that we have actually been on the same page.
The real benefit of UUIDs is the 'consistency' of the one-size-fits-most approach. If you can do without IDs humans can read out, or readable plain text logs, or compressibility, or recognisable formats for different types of ID? Then UUIDs can be used for anything from customer orders to web requests to log lines.
I wouldn't use it for assigning bank account numbers, but for most web or app stuff it's fine.