Idempotency
berkansasmaz.com
berkansasmaz.com
It does teach you how to build at least semi-reliable software because people tend to get really mad when you screw with their money. It also is a good introduction to regulation and compliance.
Idempotency was is one of the concepts you get taught on day one in that field (as kind of demonstrated by that being the example used by the author in the article).
If it had a screen and buttons, I would try to break it.
So I started striving for highly reliable systems not because there are professional bad actors out there or spammers or to achieve high performance… but because there’s another little shit out there.
I'm not sure this is the kind of failure idempotency is design to tackle.
Idempotency tackles the issue of handling multiple requests for doing the same thing, isn't it?
yeah... there should be no extra side effect if you call something more than once on it with the same parameters. ie foo(x) is the same as foo(foo(x)) is the same as foo(foo(foo(x)))
or in networking GET x is the same as GET x followed by a GET x etc... the GET request doesn't (*shouldn't) change anything.
If you start with State A and a call changes it to State B, what does running the call again do? A->B? But you’re already at B. Shit’s going to break. Redesign your system.
If there's a process with two steps, moving from Step A to B requires the process to be at Step A and it's already at step B, it should return an error. At first glance, this doesn't seem the kind of problem that idempotency is supposed to prevent...
Doing so using microservices will teach you a lot more (but may not be immediately applicable as it's not so favored).
- Useful to know when you work with APIs (as the article outlines).
- Very useful when working with background jobs. You're not gonna have a good time if those aren't idempotent.
- Good to know for interviews. I was asked to explain idempotency a handful of times, weird as that is.
I've seen this lead to some confusing back and forth between software engineers, where one engineer means one definition, and the other engineer meant the other definition.
* In the object-oriented sense, `obj.f(args...)` does not mutate `obj` after the first time.
* For pure unary functions, `f∘f === f` - that is, repeated application does not change the result, e.g. `abs(abs(x)) === abs(x)`
* For pure binary functions, `f(x, x) === x`. I'm not sure how useful this is but it seems to be used in math.
* For pure binary functions, `f(x, a) === x` and/or `f(a, x) === x`. These would probably be individually called left-idempotent and right-idempotent in some order, along with two-sided idempotent for the combination (akin to left identity (related!), left inverse, etc.); notably these are more general than (and all imply) the above. This could be extended to further arities; usually the preserved element will be either first or last in sane functions. "and not" works in one direction; many functions like "bitwise and", "bitwise or", "min", and "max" work in both directions.
Do you now reply "I've been asked this before, so your question has no effect"?
You can't check for existence to guard against duplicate creation without either locking or another approach to making the Exists-or-Create step atomic.
If you were to put in the sequence as described, you'd have something which is idempotent most of the time, until it isn't, when API requests come in and get forwarded to the endpoint before the caching server can cache the first creation.
If your API is for example: "POST /orders/create {products:[{bananas:1}], shippingAddress:{},...}", then it's not clear how you would check uniqueness, because you'd want the same person to make the same order at different times by design.
To prevent double-submit, and make it idempotent, you need to add extra information.
One way to address this is generating transaction/sequence/correlation IDs at the very start of the process. So for example when you first request the order form, a transaction / correlation ID (called in the article an Idempotency key) is generated.
Then you would include that transaction ID in the order, and could prevent duplicate use of an order ID at multiple levels of the system including the database.
One design is to have a single-source-of-truth, and have at the (internal) API layer a state machine for a transaction. Once an order has been processed then it's impossible to make another order because there's no "make order" step from the state that an order is in having been processed. Again you need some kind of correlation ID to track which state machine needs loading.
That however is less scalable than other solutions including things like event-sourced architecture, which would also prevent duplicate orders through eventual consistency although there's less (or no) guarantee in many such systems that the order submitted "first" would win the race and not be superseded by the second. Such behaviour is fine if documented.
However, if the requests of the same key (duplicates) are sent to any of several API servers, you have a problem that must be resolved as you mentioned.
In the first case, if the API server dies, the client should establish a new connection and get new idempotency keys.
It's also amazing how hard of a concept idempotent data pipelines are for some folks to understand, and how much effort they'll put into _avoiding_ idempotency.
Idempotent calls might be used redundantly to ensure that a particular state or setting exists at a particular time.
I don't really care that a particular setting was already set yesterday or an hour ago.
Conversely, it's probably a really a good idea to be sure it's on that particular setting right now, before I fire the laser!
You can embed tracking in your email to see if it's been opened, and not resend in that case, but this will be flakey as it's easy to disable email trackers.
Missing key answers like what about two requests in flight, how to store and query requests status reliably.
The article mention requests should be ACID, but the diagram has caching storage. Would redis/dynamodb work here?
https://github.com/stickfigure/blog/wiki/How-to-%28and-how-n...
The author proposes a solution that needs three steps: query a cache, execute, update the cache. These are not atomic and, therefore, not thread-safe. As a result, if a second request arrives before the first is finished, the operation will be executed again.
The diagram in the article shows “Header Idempotency Key : "A"” (honestly not sure whether the second, third and fourth of those spaces exist: the diagram’s text kerning is atrocious). This should be the header field `Idempotency-Key: "A"`, though “A” would be a bad string to use (see Security Considerations).
In a real-world scenario, this is what seems to happen to Uber Eats in India back in 2019 when everyone can gets free food because of a bug in Uber's backend.
'raise the control rods by 1 inch'
vs
'raise the control rods to a height of 4 inches'I believe idempotency is when f(f(x)) = f(x)
You can abuse notation slightly to see this as f^2(x) = f(x). (This kind of notation is actually used enough to be comfortable to most mathematicians.) From here it's not much of a stretch to eta-contract this to f^2 = f.
And so we've ended up back in the same place...
Naming something also helps us to think about it. "A square" is also an obvious concept, but by naming it you can reason about it more easily, and use it to define more abstract concepts later. One of the main reasons we - humans - dominate the earth, is because we evolved an advanced language, which allows us to name ideas and build upon them.
The software industry is littered with these words. Design patterns especially. Rob Pike once referred to it as infatuation with nomenclature.
That's where I stopped reading.
Now imagine you keep getting interrupted in the middle, but really want to read it all, thus gradually filling the whole of your brain capacity.
Uhh thanks chatgpt.