But what I don't understand is how this wasn't caught after the first failure? Does this company not have any logging? Shouldn't the fact the backend is attempting to reuse UUIDs be immediately obvious from observing the error?
But what I don't understand is how this wasn't caught after the first failure? Does this company not have any logging? Shouldn't the fact the backend is attempting to reuse UUIDs be immediately obvious from observing the error?
UUIDs are just 128-bit values. They might be conventionally encoded for humans as hex, but storing them as 36-byte (plus a few more for length) strings is a pointless waste of both space and performance.
I don't think 128 bits vs 36 byte performance it's a main concern right now
36B vs 16B today, tomorrow you need an array of it, and now it isn't cache aligned, and more than twice the overhead.
Most likely instead of manipulating a 36B fixed length string, it is handled as a dynamic string, for extra runtime memory allocations, most likely consuming at least 64B per allocation. Etc etc.
Do this all over the codebase and now you know why all the moderne software is a sloth on what was a supercomputer 30y ago.
It’s also not just the size itself. Despite being fixed-size in practice, these are variable-sized strings in application code which now means gajillions of pointless allocations and indirection for everything. There are a ton of knock-on performance consequences here, all on the most heavily-used columns in your data model.
Worst of all should they actually succeed, this is going to be absolutely excruciating to fix.
But in either case – MySQL or Postgres – they’ve still made the classic mistake of using a UUID as a PK, which will tank performance one way or another. They’ll notice right around when it starts to matter.
I guarantee you that they _will_ have another production bug like this sometime in the future (every fast paced project will). You'd hope this next one wont take 5 days to identify.
I am also very confused about the apparent lack of logging or recourse to logging. It's been a while, but if I recall correctly ECS should automatically propagate the resulting Duplicate Key exceptions which were presumably occurring to CloudWatch without a bunch of additional configuration - was that not happening? If it was happening, did no one think to go check what types of Exceptions were happening overnight?
Specifically asking why did it take so long to detect and why did it take so long to diagnose is useful in these situations.
Type 1 tries to find the error message and figure out what it really, really means by breaking down the error message and system.
Type 2 does trial and error on random related things until the problem goes away.
I hate to say that I've seen way more type 2s engineers than type 1, but maybe I’m working at the wrong companies.
Here we are talking 1.65 MILLION CAD $ backed YC company