UUID v7
commitfest.postgresql.org
commitfest.postgresql.org
I don't know how to navigate commitfest (nor would I probably understand the source code to begin with), but the reason I'm confused is that you can already use all the proposed draft UUID implementations in Postgres (as long as you generate the ID application-side). In fact, PG will happily accept any 128 bit ID to be inserted into a UUID column, as long as it's hex encoded – even the dashes are optional.
It's really more about convenience, and maybe a bit of speed.
1: https://gist.github.com/fabiolimace/515a0440e3e40efeb234e126...
If you want to do something smart like encoding your node ID within the value, or prefixing a timestamp for sortability, then sure, do that in your application. No one else really needs to care how you produced your 16 bytes. Just do some napkin math to make sure you're keeping sufficient entropy
I'm not sure "UUID" even needed to be a column type, versus a "INT16" and some string formatting/parsing functions for the conventional representation (should you choose to use that in your application). You could also put IPv6 addresses in the same type. Though I guess this depends on how much you think the database should encode intention versus raw storage in types
UUID isn't about how it's done, it's about what it is.
Instead of everyone doing something differently, everyone can just comply with UUID.
Instead of having to repeat it across your docs that the IDs of this entity are sortable, you can just say they are UUIDv7. If someone wants to extract the timestamp from your ID, they don't need to figure out which (32, 48, 50?) bits are the timestamp nor what resolution the timestamp has because you can tell them UUIDv7.
You don't have to write your own validation functions because you can tell the database that this hex string is a UUID and it can do it for you.
You're probably making the case for Sqlite here which is very minimal, but something more full-featured like Postgres, I prefer these conveniences. I can tell because whenever I use Sqlite in a case where I could've used Postgres, I regret it!
Most functional things related to e.g. embedding the record creation time within the ID is one of those "that's cool, but I've never seen anyone do it" kind of things. If you need to sort records by when they were created, there are probably three or four happened_at fields on the record you'd use (created_at in this case). If you need the exact time; those are there for that.
Counter-argument: Well, you can save a few bytes on every record by getting rid of the created_at field and just using a UUIDv7. Maybe, but I've never seen anyone do it. What if you need to change the time the record was created? Are you planning to explain to all your integration providers the process of extracting a timestamp from a UUIDv7? What if you need to run complex SQL timestamp functions on created_at? Etc. Its cool, but it never actually happens.
Once we enter the domain of "using the node id or timestamp or something to reduce the probability of ID collision", that's a totally reasonable responsibility within an ID's set of concerns. But, that's a very different need.
> You don't have to write your own validation functions
Why are we validating IDs?
> but something more full-featured like Postgres, I prefer these conveniences.
Agreed. I am a vocal UUID hyper-hater. UUIDs should be destroyed, and humanity would be (oh so slightly) better off if they had never existed. But, they're still a thing, and I think its cool that databases have hyper-specific types like this.
My wish is Postgres would have other more sane automatic ID types and gen capability, in addition to uuid & autoincrement.
Random primary keys are bad, but exposing incremental indexes to the public is also bad, and hacking on a separate unique UUID for public use is also bad. UUIDs are over-engineered for historical reasons, and UUIDv7 as raw 128 bits without the version encoding would be nicer.
But, to the end-user it's just a few lost bits in a 128-bit ID with an odd standard for hyphenation. The standardization means you know what to expect as developer, instead of every DB rolling their own unique 128-bit ID system with its own guarantees and weirdnesses.
Is it bad just because of the extra bytes used, or something else?
I suppose the index order may still be relevant for PostgreSQL?
The realistic answer is: it isn't, because pre-UUIDv7 there was literally nothing about the UUID spec that conferred more capability than just a random string. And, truly; people used them as "just gimme a random string" all the flipping time. The pipes of the internet are filled with JSON that contains UUIDs-in-a-string-field, 4 bytes wasted to hyphens, 1 byte wasted to a version number, none of that is in service to anyone or anything.
People have reason to do those things, and oh boy do you want to know that it's happening. With UUID, over-engineered as it may be, you know what you're asking for and can see what you're getting - truly random, server namespaced, or chronologically sortable.
2. Being upset over 4 bytes wasted to hyphens but not being upset about JSON itself seems hypocritical. JSON is extremely wasteful on the wire, and if you switch to something more efficient you also get to just send the UUID as 16 bytes. That's a lot more than 4 bytes saved.
Over JSON you can still base64 encode the UUID if it's not meant to be user-facing.
The version/variant bits are the pointless part. Of course if you put the 16 bytes on the wire you would still have some encoding (perhaps 22 base64 characters?) that requires decoding/validation, but in memory and in your DB it's just 16 bytes of opaque data.
With which UUID ? UUID v1 ? UUID v2 ? UUID v3 ? ... UUID v7 ?
version from Latin vertere "to turn, turn back, be turned; convert, transform, translate; be changed"
variant from Latin variare "change, alter, make different,"
In my own UUID/GUID code I have taken to calling them "behavior" and "layout", respectively: https://github.com/okeeblow/DistorteD/blob/NEW%E2%80%85SENSA...
4.1.3 The version number is in the most significant 4 bits of the time stamp (bits 4 through 7 of the time_hi_and_version field). The following table lists the currently-defined versions for this UUID variant. The version is more accurately a sub-type; again, we retain the term for compatibility.
It's recognized in the RFC and all you've done is broke compatibility for fashion.
The practical necessity of mixing UUID versions, along with other 128-bit UUID-like values, means that the collision probabilities are far higher in many non-trivial systems than in the ideal case of having a single type of 128-bit identifier. There is a whole separate element of UUID-like type engineering that happens around trying to mitigate collision probabilities when using 128-bit identifiers from different sources, some of which you may not control.
Having 128-bits is the only common thread across these identifiers which everyone seems to agree on.
If you're building for a single application or data type, sure do your thing, have at it. If you're trying to coordinate UUID spaces and generation across thousands of different applications and data types, like large data pipelines, then this matters a lot.
Also, having native database support (like indexing, filtering, etc.) improves efficiency for these types of workloads.
Conflating timestamps and uniqueness is a conflation. These concerns are best separate. If you put a timestamp in your UUID, your UUID now has a timestamp. You can't remove it if you don't want the timestamp to be a part of the uniqueness any more.
- sortable by insertion
- less vulnerable to sequence prediction attack
- allows partitioning of tables in the future
Sequence prediction attack is a problem when you want to use identifiers in a public API. For example, a user or competitor can iterate through your product catalogue by incrementing the key. Sorting by inserting is a common use case as well. You can achieve this with an auto-incrementing primary key, however, this will create issues when you need to partition the table. Of course, you may never need this functionality, but there is a reason UUID7 have been added, as they are very useful in certain scenarios.Which leaves partitioning, something that very few applications will ever need, even if plenty of developers hope they will. A slightly less weird use case, but perfectly doable up to a point using regular sequences.
I see a solution looking for problems.
Then select * from item where id=123 and vf='ixkdS1Xb'
I fully agree and support that auto-incrementing integers should be used whenever possible. My preference for UUIDv7 over UUIDv4 is solely that they’re less likely to wreak havoc on the DB, if devs insist on having a UUID PK.
Sadly this is very much a thing that ORMs do all the time.
|------------------------------|-----------------------------|------------|----
| primaryKey (bigInt-internal) | publicKey (uuidv4-external) | created_at | ...
|------------------------------|-----------------------------|------------|----
(2) |------------------------------|-----------------------------|------------|----
| primaryKey (uuidv4-internal) | publicKey (uuidv4-external) | created_at | ...
|------------------------------|-----------------------------|------------|----
(3) |------------------------------|-----------------------------|----
| primaryKey (uuidv7-internal) | publicKey (uuidv4-external) | ...
|------------------------------|-----------------------------|----
(4) (not recommended)
|---------------------|----
| primaryKey (uuidv7) | ...
|---------------------|----
---- (1) [X] sortable by insertion, [X] timestamp
(2) All of (1) and [X] transferable between databases
(3) Use UUIDv7 as a primary key for internal and UUIDv4 for external. App or SELECT statement will need to extract the timestamp from UUIDv7 if you need to use it. Also, if you're using a DB Client you can't just view the 'created_at' column to get an idea of when a row was created.
(4) Use UUIDv7 as a primary key for internal & external use.Over time conflicts between the id's primary job (uniquely identifying something) and the extra semantics can arise, and the solutions tend to get pretty messy.
Here we have a unique id that embeds a timestamp. The classic conflict here is with privacy/security. A UUIDv7 user id tells you when the user was created. A UUIDv7 of a medical record tells you when some medical event occurred.
There are things whose identity is inherently time-based and not private, so I'm not giving a blanket recommendation to not use these. Just understand what you are signing up for.
For a database, you can use bigints for primary ids but only internally. Then you also have an external random (v4) uuid... and a timestamp if you want, for that matter -- now that it's a separate column, you can expose/hide it on a case-by-case basis, depending on need. So this gets you the benefits of a uuidv7 but maintains flexibility, though at the cost of some complexity and extra bytes/record.
Other conflicts can arise too, and they can be hard to always foresee, so generally be careful about extra semantics in unique ids.
If that's an issue for you, you can get around this in a variety of ways, as they mention: you could use an associative table that maps the externally-exposed random ID to an internal-only ID.
Our competitor was able to ascertain, based on that timestamp, when our customers contract was up and was (briefly) able to poach some customers by underbidding us until we corrected this.
I'm so naively the "take the high road" kind of guy, that I just assume everyone (or every company) should just do the right thing. Stealing customers in this way from a competitor, I have no ability to rationalize such an action.
And this kind of makes me scared, like if I were to ever own a business, I just know I'm swimming with sharks with no ability to defend.
I believe your story, but it's just crazy to me. Go earn a customer's business in a legit way, not be stealing data from a competitor.
Anyway, the arc of the universe bends toward justice: this (former) competitor got sold for parts a few years later.
"locationId": "53c24146ef0b601b77974fcd"
They took the first four bytes (53c24146) which is a timestamp that represents 1405239622 seconds since Unix epoch. Our website clearly stated we work off annual contracts (a norm for our industry) - it wasn't secret information. So from this timestamp they could ballpark when a customer's contract was up.IMHO, leaking an ID always impose a risk. It is always have negative impact if leaked, no matter how perfect your system is.
While it’s no longer on ranks on the top 10 web vulnerabilities, gaining internal insight to systems is one of first things you do when infiltrating.
But people are messy and lazy. Nowadays, you ask for GDPR data and people give you CSVs with all their real table and column names.
Sometimes when you are just a little inside, figuring out an id is like figuring out a password (particularly with uuid as opposed to a sequence). Real nice if it leaks easily.
We must live in a different universe. I'd wager to say that over 90% of all backends leak their primary key when speaking to the front-facing client.
That sounds like a problem which should be solved by making database engines not assume keys have some sane ordering, not by putting timestamps in UUIDs.
It is quite possible to do cluster-style indexing on UUIDs through disk on a single server at rates of tens of millions per second, I do it every day, just not with your typical ordered-tree architectures. Many popular database engines are not designed to make this particular scenario perform well.
It tells you when the event was documented. If the event didn't contain a date time stamp itself, I would be highly surprised, because what other value is there in documenting it?
The security problem here is inherent in the practice and your choice of primary key isn't a material factor at all.
Do you imagine there's a public CRUD database with simple Rails style accessors that can drill all the way down to individual event records inside my health information? And that, somehow the leak of a primary key in a URL might give away the fact that _something_ happened to me, medically, 12 days ago?
Just as a side note: it may also leak the location, not just the time. An that is enough e.g. for disproving an alibi or leaking an important commercial secret (if you are in the same location as competitor HQ, for example).
I cannot rightly apprehend the scenario you are describing. A medical provider might generate a UUIDv1 and add it to a record of mine, and this will somehow destroy my ability to have an alibi in court with respect to corporate espionage?
I'm not in a bond movie, I just need to keep track of events and have them sort in chronological order
For me the main challenge was that it's still considered a draft (AFAIK). It may be unlikely to change, but if it does I'd rather not have to deal with persistent UUIDv7 data generated per some previous spec.
Also, if I really want/need UUIDv7, it's not that hard to create an extension that generates UUID in arbitrary ways, including the proposed v7.
there's already well maintained extension https://github.com/fboulnois/pg_uuidv7
It's slightly different from recommendations by draft RFC version (there's no counter), but fully within spec requirements. From practical point there's no difference at all.
It would be quite unfortunate to end up with a UUID v7 in PostgreSQL that’s not quite the standardized one because the patch got merged too quickly.
EDIT: here is the IETF working group page https://datatracker.ietf.org/wg/uuidrev/about/
Their milestone seems to submit the final proposal by March.
The chances of that seem extremely low at this point. The contents of a version 7 UUID have not changed since work started on RFC 4122 bis in October 2022: https://author-tools.ietf.org/iddiff?url1=draft-ietf-uuidrev...
There's clearly little pressure to rush this, considering it's not difficult to add a custom function generating UUIDv7 ...
Correct. The sortable nature of UUIDv7 improves database performance and index locality by helping the index be more efficient since rows are inserted in a predictable order instead of scattered randomly.
UUIDv7 spec: https://www.ietf.org/archive/id/draft-peabody-dispatch-new-u...
The optimization relevant to v6 also applies to v7. The difference is that v7 UUIDs are not directly compatible with v1 UUIDs.
The current draft can be found here: https://datatracker.ietf.org/doc/draft-ietf-uuidrev-rfc4122b...
It seems the algorithm is equivalent to:
function nanoid(alphabet="ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-_") {
return Array.from(crypto.getRandomValues(new Uint8Array(21)), i => alphabet.charAt(i & 63)).join('');
}
(Maybe String.fromCharCode() is slightly faster than Array.join(), but I doubt it matters much.)On node.js it's even easier, since url-safe base-64 encoding is supported natively:
Buffer.from(crypto.getRandomValues(new Uint8Array(16))).toString('base64url').substring(0, 21)
Do we really need an entire Github project dedicated to a 1-liner?* Using uncoordinated random generation in a large address space for IDs is often a very useful idea
* 128 bits is a good rule of thumb for "will never ever collide"
* Making this into a heavyweight "UUID" concept, with it's own bespoke string format, and 7 different standard ways to generate them, feels like a ridiculous waste of cognitive effort that makes such a simple concept appear opaque and magic. If you want to encode other data (timestamp, node ID) in 16 bytes you can still do that of course. There's no need for anyone else to even know. Just do some quick calculations to ensure you haven't eliminated too much entropy
It depends.
True random number are slow. Fast PRNG are prone to collision. That's why we have different specs and versions of UUIDs.
Not in this decade. You can slam out a million securely random 128 bit numbers per second per core. For numbers you will store, the effort to store them is orders of magnitude greater than the effort to generate securely.
That would mean it's easy to update old apps, as we would only have to change the default value for the columns.
It’s a good thing to keep in mind though for sure.
I don't have a resource of the top of my head to present to you, but in the least keyset pagination is superior to the offset one because it does not get invalidated by new inserts.
You cursor on timestamps, or serial numbers, or etc., not IDs.
Integers don't scale because you need a central server to keep track of the next integer in the sequence. UUIDs and other random IDs can be generated distributed. Many examples, but the first one that comes to mind is Twitter writing their own custom UUID implementation to scale tweets [0]
[0]: https://blog.twitter.com/engineering/en_us/a/2010/announcing...
This was nice because if the script failed half way through I could easily lookup which ids were already imported and continue where I left off.
The point is, this property of UUIDs occasionally comes in handy and it’s a life saver.
postgres=# CREATE TABLE foo(id INT, bar TEXT);
CREATE TABLE
postgres=# INSERT INTO foo (id, bar) VALUES (1, 'Hello, world');
INSERT 0 1
postgres=# ALTER TABLE foo ALTER id SET NOT NULL, ALTER id ADD GENERATED
ALWAYS AS IDENTITY (START WITH 2);
ALTER TABLE
postgres=# INSERT INTO foo (bar) VALUES ('ACK');
INSERT 0 1
postgres=# TABLE foo;
id | bar
----+--------------
1 | Hello, world
2 | ACK
(2 rows)I’m sure there’s a way to get it to work with integer ids but it would have been a pain. With UUID’s it was very simple to generate.
> I’m sure there’s a way to get it to work with integer ids but it would have been a pain. With UUID’s it was very simple to generate.
IME, if something is easy with RDBMS in prod, it usually means you’re paying for it later. This is definitely the case with UUIDv4 PKs.
But I get you like integers so whatever works for you, I just don't think they're the right tradeoff for most projects.
They most assuredly do scale. [0]
Also, Slack is built on MySQL + Vitess [1], the same system behind PlanetScale, which internally uses integer IDs [2].
[0]: https://www.enterprisedb.com/docs/pgd/latest/sequences/#glob...
[1]: https://slack.engineering/scaling-datastores-at-slack-with-v...
[2]: https://github.com/planetscale/discussion/discussions/366
If you're using MySQL maybe integer ids make sense, because it scales differently than PostgreSQL.
If you're referring to not having conflicts between distributed nodes, that's a solved problem as well – distribute chunked ranges to each node of N size.
If you can't manage minor levels of coordination because your database is on fire, the problem is that your database is on fire.
In general you shouldn't need to make a roundtrip to produce an ID.
The distributed database needs a coordination system anyway, so it's not an additional point.
> In general you shouldn't need to make a roundtrip to produce an ID.
Did you forget the context over the last week? We're already talking about reserving big chunks to remove the need to make a roundtrip to produce an ID. There would instead be something like one roundtrip per million IDs.
Nope! Distributed databases do not necessarily need a "coordination system" in this sense. Most wide-scale distributed databases actually cannot rely on this kind of coordination.
> Did you forget the context over the last week? We're already talking about reserving big chunks to remove the need to make a roundtrip to produce an ID. There would instead be something like one roundtrip per million IDs.
OK, it's very clear that you're speaking from a context which is a very narrow subset of distributed systems as a whole. That's fine, just please understand your experience isn't broadly representative.
I'm assuming a system that tracks nodes and checks for quorum(s), because if you let isolated servers be authoritative then your data integrity goes to hell. If you have that system, you can use it for low-bandwidth coordinated decisions like reserving blocks of ids.
Am I wrong to think that most distributed databases have systems like that?
> OK, it's very clear that you're speaking from a context which is a very narrow subset of distributed systems as a whole. That's fine, just please understand your experience isn't broadly representative.
Sure, but the first thing you said in this conversation was "Whatever is distributing the chunks is still a point of central coordination." which is equally narrow, so I wasn't expecting you to suddenly broaden when I asked why that mattered.
Not sure why.
> because if you let isolated servers be authoritative then your data integrity goes to hell
Many AP systems maintain data integrity without central authorities or quorums for data.
> Am I wrong to think that most distributed databases have systems like that?
No, not wrong! Just that it's one class of distributed systems, among many.
It reminds me a bit of the microservices trend. People tried to mimic big tech companies but the community slowly realized that it’s not necessary for most companies and adds a lot of complexity.
I’ve worked at a variety of companies from small to medium-large and I can’t remember a single instance where we wish we used integer ids. It’s always been the opposite where we have to work around conflicts and auto incrementing.
I've personally ran MySQL in RDS on a mid-level instance, nowhere near close to maxing out RAM or IOPS, and it handled 120K QPS just fine. Notably, this was with a lot of UUIDv4 PKs.
I'd wager with intelligent schema design, good queries, and careful tuning, you could surpass 1 million QPS on a single instance.
- client-side generation (e.g. can reduce complexity when doing complex creation of data on the client side, and then some time later actually inserting it into to your db)
- sequential ids leak competitive information: https://en.wikipedia.org/wiki/German_tank_problem
- Global identification (being able to look up an unknown thing by just an id - very useful in log searching / admin dashboards / customer support tools)
Those two sucks for us right now (planning to move to UUIDs).
Wouldn't Snowflake IDs also solve that problem? A Snowflake ID will fit within a signed 64-bit int.
https://en.wikipedia.org/wiki/Snowflake_ID
The nice thing about a Snowflake ID is that you can encode it into 11 characters in base 62. If I have a UUID, I'm going to need 22 characters. Maybe that doesn't really matter given that 11 characters isn't something someone will want to be typing anyway and Snowflake IDs do require a bit of extra caution to make sure you don't get collisions (since the number you can make per second is limited to how big your sequence generation is).
If your system ever becomes distributed you will sing the praises of whoever choose UUID over an int ID, and if it never becomes distributed UUID won't hurt you.
Note: this is for web systems. If it's embedded systems then the overhead starts to matter and the usefulness of UUID is probably nil.
Less of an issue if you have total control of the operational environment and code base, but that is not always the case.
The main reason non-probabilistic UUID-like types are used for high-reliability environments is that it is easy to verify the correctness of the operational implementation. It isn't that difficult to deterministically generate globally unique keys in a distributed system unless you have extremely unusual requirements.
I’ll again point out (I said this elsewhere in a post today on UUIDs) that PlanetScale uses int PKs internally. [0] That is a MASSIVE distributed system, working flawlessly with integers as keys. They absolutely can scale, it just requires more thoughtful data modeling and queries.
[0]: https://github.com/planetscale/discussion/discussions/366
Some comments: https://news.ycombinator.com/item?id=39261469
https://news.ycombinator.com/item?id=36433481
And have a look at this post from last year: https://news.ycombinator.com/item?id=36438367
EDIT:
For those who misunderstood: I'm very much pro-UUID, and against serial autoincrement server-generated IDs. With the exception when you need to heavily optimize for speed and/or index storage space. And even then there are hybrid solutions like using UUID externally, and serial IDs internally.
Generating IDs away from the database is mainly done due to design restrictions demanding it, not because it is beneficial.
But client-generated IDs make idempotency easier and remove whole classes of errors. They're typically a huge win.
2. the UUID by itself doesn't authenicate or authorize anything
3. there is a small chance of collision, and it can be handled on the backend/persistance/DB layer, i.e. return error to the client in case of collision and ask to generate a new UUID
4. many non-trivial and/or CQRS/ES apps work like this
5. if you are really paranoid, you can push down UUID generation logic to BFF (Backend-For-Frontend) layer
6. Lots of DistSys problems can be solved with client-generated IDs. But most people mistakenly think that DistSys applies to backend only, and exclude clients from the picture.
7. Scaling RDBMS (especially Postgres) is hard. UUID generation is slower than serial bigint, so it's best to keep it outside DB layer fo this reason also.
8. Client-generated UUIDs help to make client requests idempotent, and enable error handling with retries (although it's better to add an additional layer of request idempotency with IDEMPOTENCY_KEY HTTP header, or GraphQL Relay's clientMutationID).
the malicious actor will need to somehow get the authorized session first, then get the real UUID which already inserted into the DB.
And now she will be able to do DDoS by replaying create_entity requests with already existing UUID.
But ... the same scheme equally applies to server-generated IDs also when used for updates instead of insert/create (even if there is a translation layer between internal and external IDs) ;)
In my implementations this create_entity will also have an idempotency key, so after first successful request, the replayed/retried requests will hit the cache only, and skip the database.
In case the attacker will also change idempotency_key for each replayed request, then the only remedy is to monitor for this scenarios, or reject requests with the same {sessionID, requestName, enitityUUID} and different idempotencyKey-s over short time intervals.
But again, this equally applies to any type of server-generated ID.
I'm not saying this is a big issue, but saying "P_collision = 1 / 2^122" is misleading and gives a false sense of security. P_collision is 100% if a malicious user can specify any id they want and wants to specify a colliding one.
I don't understand why you're bringing in idempotency keys. The way to fix this is to reject client generated ids that already exist.
This is bad UX and bad design, usually human-readable event slug will be shown in the URL.
But I agreed that since UUIDs are not encrypted, they can be acquired by the attacker (but this equally applies to any other unencrypted data handled by the client).
> Suppose there's also an api where you POST a JSON body to /api/v1/event to create an event. This uses a client generated ID so the body contains the title, location, etc and the supposedly newly generated id. A malicious client could submit an existing uuid instead of generating a new one, and the server would need to reject this.
Again this is bad design, and the attacker will also need to either hijack the session, or to steal JWT token.
> I'm not saying this is a big issue, but saying "P_collision = 1 / 2^122" is misleading and gives a false sense of security. P_collision is 100% if a malicious user can specify any id they want and wants to specify a colliding one.
I agreed with you, but the attacker can copy any ID, including server-generated ones. So I don't understand the problem. Any external data need to be validated be it client-generated or server-generated. Do you claim that somehow validating UUIDs for uniqueness in the DB layer is more expensive than any other validation of the external data?
> I don't understand why you're bringing in idempotency keys. The way to fix this is to reject client generated ids that already exist.
Please read again, yes all external data need to be validated, so it equally applies to server-generated IDs (only they will need to be validated on UPDATE-s and DELETE-s instead of INSERT-s).
Idempotency keys just reduce the load on the hard-to-scale RDBMS database layer, so it will not be hit on every retry or replayed malicious requests.
Sure you are. You're trusting the client to send you valid data.
> 2. the UUID by itself doesn't authenicate or authorize anything
Okay, better hope no one ever considers the UUID to be a unique randomly generated token.
Besides that, users might be tempted to submit "vanity" UUIDs if they get to decide their own identifiers, breaking assumptions about the system.
> 4. many non-trivial and/or CQRS/ES apps work like this
Cool, if your friends jump off a bridge, you gonna follow them?
Reason: I'm not a machine, so I dislike generating UUIDs by myself ;)
But yes, it's not a true semantic key since the first name usually doesn't come from an immutable properties of the child, but the last name can be a semantic key, derived from the properties of the family (e.g. geographic origin, profession, hair color, look, etc.)