Idempotence now prevents pain later
ericlathrop.com
ericlathrop.com
"Need to send a notification email when x condition becomes true". Naive way: during processing, check the condition and call the SendEmail() function. Idempotent way: Run a query that finds all x conditions, join to a list of notifications based on email+id+time, and only if there's no entry, send the notification and save it to the list.
"Send a report at midnight". Naive way: have a cron entry that runs at midnight or run a loop that checks the time, and if midnight, sends the report. Idempotent way: cross-check against a list of reports that should have been sent, and if missing, send it. Run this check every few minutes.
The nice thing about this is it doesn't stop you from also doing the event-based stuff (for "real-time" notifications, for example), so long as it's the secondary approach from design point of view. Ironically when everything is working fine, the event-driven stuff will actually be doing all the work -- but as soon as there's a failure (like, the system happens to reboot at 23:59) you'll be grateful to build this way.
I could write paragraphs about the nuance, different approaches, failure modes and the trade-offs with all this -- suffice to say treating events as optional in event-based systems has served me well.
That’s to say, this comment probably just saved or made me money. I’ve some hands on but clearly not as much as you exhibit. Thanks.
Target audience: Sr SE, reporting to next level up and perhaps offering tradesoffs, or advising peers or reports working on those systems and need a refresher to give a complete answer to a question.
How can I find more wisdom like this?
By reading more Hacker News ;-)
Seriously though, if you want learn a different way of building applications where loads of these things pop up, look into Actor Model. Even if you are not going to use Actor Model, reading up about it will teach you a lot about "realtime" applications... My main epiphany was that realtime applications do not exist. Rather they cannot exist. Because our reality is not realtime (or our perception of it is always delayed because our brains/awareness is slow) - we can never truly capture the present moment. Data we see is always in the past. Reality is eventually consistent (from our perspective). All things in our reality runs concurrently, stateless and in their own little time bubble. You can also read up on the concepts behind Event Sourcing.
Anyway, it is a huge rabbit hole, but reading up on the ideas behind Actor Model might just blow your mind. I'm convinced it will make a comeback because it is the natural way our reality works and Actor Model mimics it the closest, and it will also work better for multi-core cpu's. For non-tech reading, try The Power of Now by Eckhart Tolle. Then try to reflect on (relational databases) + (actor model) + (event sourcing) + (power of now) + (brain in a vat philosophy) + (try to to figure out at what speed the universe "renders", how fast our brains can think, how much delay there is, what limits do our body's sensors limits impose) + (try to figure out how closely Object Oriented programming can mimic reality (hint: all programming are tiny simulations that try to mimic some tiny part of our reality)) + (read up on the history of time keeping and the huge rabbit hole of calendars, and how most programming languages attempts to handle time (badly for the most part, it's a legit rabbit hole on it's own)).
After a year or two of hard reading, reflection and teachings (and wrecking your head), you may have good insight into the nature of programming and the fallacies of trying to build realtime systems, or trying to mimic our reality in a computer in general. It is insanely complex, we as programmers merely touch a grain of sand of the beach of complexity. We have no idea how simple our most convoluted systems are in the face of nature, yet we convince ourselves that we KNOW how things work and that we are in control at all times.
Good luck!
Most of us think our applications are realtime (never mind the OS schduling cpu time for your code, so inherently not realtime) but this is only because it behaves as expected, while there is not too much load nor too much data. The moment you start having significant computational load or start having too much data (aka the cpu, memory or storage starts choking), or when you start to split data apart (aka distributed applications) then the problems with "realtime" applications starts to become obvious.
If something goes wrong, to fix the immediate problems would be to retry/rerun the simulation, but that might result in duplicating your intents, which results in the simulation and its state as invalid. Then you either (hopefully) have some mechanism to de-duplicate whatever happened more than once or you have to start you simulation from scratch after fixing the data manually.
Or just design the core of the system with Idempotency in-mind from the start so you have a good way to handle such scenarios. Embrace that most things can/should be handled in a eventually-consistent way. Very few things in life has to happen RIGHT NOW (good luck with that).
An interesting spin on realtime systems are computer games, where we refresh a view of the simulation every few milliseconds (aka 60fps). Every frame gets rendered based on the inputs (keyboard/mouse)from just-after the previous frame. So in effect, the current frame being drawn is just a reaction based on previous inputs. So let me ask you this: in a game engine like this, where would you assign the value of "now"? Right after the previous frame? When capturing user input? When rendering the new frame starts? When showing the new frame but before capturing the new inputs? Where is reality's "now" moment in a computer game? Let's ignore the delays between keyboard/mouse input and drawing from CPU/GPU to the physical screen and lets assume this is a single player FPS and not a multiplayer turn-based game. And lets ignore sound processing too - just think of the simplest of game loops - should we consider them as realtime? Ponder that.
By the time you perceive anything, and even more so by the time you react to it, it's already in the past.
There is no "now", it's all an approximation.
My younger self did that, quite a lot, thank you. I never felt comfortable with various approach of intermixing input handling, physics/game logic and rendering - whether to run them sequentially, decouple some or all of them, whether the physics should be "look-ahead" or "look-behind"... I ended up picking the pattern that leads to the most stable and deterministic behavior (look-behind physics in lockstep with input and on fixed update rate, rendering independent and done variable-rate), and stopped thinking about it.
So thanks for mentioning it in this context, it makes me more comfortable to hear from someone else that "now" in games is ill-defined on philosophical level, and I shouldn't have worried about it as much as I did.
Just keep a log of all the events that have already happened (crucial - I have seen "event logs" that were more like "request logs") and to find out the current state of the system - if the notification or report has has been sent - by examining this log.
The lists you are describing kind of sound like an event log of sorts.
It seems to me that this distinction between primary and secondary approach is not really neccessary.
Describe the final end/desired state and let the machine figure out how to get there, progressing through the work until it's caught up even across restarts or other boundaries.
One of the experiences I had was applying it to provisioning with idempotent shell scripts, combined with a trivial tool like https://github.com/lloeki/apply
Creating a VM in DO would add the basic SSH keys, then just run the thing against the new IP. Boom, VM ready. Want to make existing VMs config up to date? Apply. Boom.
The whole thing ridiculously scaled up in a totally unexpected way, saving hundreds of hours of headaches and mistakes.
Why not ansible/puppet? Check out the rationale in the readme.
This was common when I was writing puppet. It was easy to have your entire puppet run take minutes - and want to run every few minutes - while consuming a LOT of cpu time including on a centralised server and not just the end host.
There are of course ways and strategies to this and not every situational falls afoul of it. But it is something you very much need to be aware of.
I otherwise whole heartedly endorse this.
Definitely! And on top of that, you probably want to apply some business logic to what you're doing.
To expand on the example of "Report should run at midnight":
If you run the check regularly, you can constrain the date range you're checking to just since the last time you checked, and store this date either in memory or persisted depending on your use case, technology, and how you want system failures/restarts handled.
Similarly, if the system was down for several days, it might not be relevant to send all missing reports, but instead just the last one.
On the other hand, if "daily reports" are required for auditing purposes, it probably does make sense to send all of them. But you have to be careful so on the very first run it doesn't go "Hey, there are 18,726 missing reports since Jan 1, 1970!"
It's also important to write the report itself to be idempotent. You can't write it so it reports to "current" time with the assumption it runs at exactly midnight. Instead you have to actually constrain the dates to exactly what you want. Even aside from the idempotent invocation I'm talking about, as your system gets big you'll get thousands of reports that all need to run at "midnight" and that just isn't practical: maybe they run all sequentially, and some will not actually execute until several minutes after midnight.
instead you simply see if there's a report for yesterday, and if there isn't make one.
the edge cases are still there, see replies
Ha! You are biased toward the word "Idempotent way". Above - you have totally different steps under "Idempotent way". So, your process is different for each of them. There's no tech differentiator and it should be - as that's how you implement "Idempotent way"
I am not sure how to perform that rewording/transformation in general though or even for the examples above.
However, that description relies on implicit state: it assumes a world where all dormant customers have not been charged. As we all know, relying on hidden implicit state is the source of all evil. When you describe your problem using verbs, you are essentially modeling it in terms of a diff between how the world is and how you want it to be. But that diff is only correct if you have a perfectly correct description of how the world currently is.
It's better to take a step back and think of your problem purely in terms of the world you want to end up in as a result: "all dormant customers have been charged". Framing it that way makes questions like "which dormant customers have already been charged?" more obvious to consider.
Great framing!
Isn't this the difference between imperative and declarative statements/code/systems?
And then the question adds state, "as of yyyy/mm/dd, which customers were more than $30 into credit?"
To give another simple example as the OP - Suppose you have a product that relies on time series data. For demo purposes you might create a curated data set to present to clients, but the presenter doesn't want to show data from 2019 as the "most recent"
Naturally, you decide to write a script. Do you
A) Write as script that moves the data forward by 1 week explicitly, and simply run this once per week or
B) Write a script that compares the current date to the data and moves it forward as much as it needs
At first glance, these two approaches work the same, but what if (A) triggers twice? What if it runs once every 6 days by mistake? (B) is idempotent however - subsequent executions won't change the state. It's usually impossible to predict all of the ways that software breaks, but designing with idempotency in mind eliminates a lot of them.
An idempotent change would be to pass in the current time instead of checking system time. In this case, as long as the input is the same, the result is the same. You could use cached results, but most likely you want to use new inputs.
What Is Idempotence? - https://news.ycombinator.com/item?id=19570815 - April 2019 (51 comments)
Idempotence: What is it and why should I care? - https://news.ycombinator.com/item?id=17804617 - Aug 2018 (73 comments)
You know how HTTP GET requests are meant to be idempotent? - https://news.ycombinator.com/item?id=16964907 - May 2018 (304 comments)
Implementing Stripe-Like Idempotency Keys in Postgres - https://news.ycombinator.com/item?id=15569478 - Oct 2017 (41 comments)
APIs, robustness, and idempotency - https://news.ycombinator.com/item?id=13707681 - Feb 2017 (50 comments)
A simple distributed algorithm for small idempotent information - https://news.ycombinator.com/item?id=7276491 - Feb 2014 (14 comments)
Idempotent Web APIs: What benefit do I get? - https://news.ycombinator.com/item?id=5662138 - May 2013 (53 comments)
A word like that is particularly easy to search for:
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
That’s when I came across Qmail and studied it’s design. Qmail was designed so well that it literally absolutely would not send an email twice.
But my scheduler did.
To send a few hundred thousand emails more than once is not good is not good for your reputation lol.
It taught me a valuable lesson about Idempotency from an impressionable stage of my career and now, with creating workers running under Sidekiq, Oban and other job schedulers that run very frequently, this has become quite useful.
At one point I used to lean on Redis to “help” me be more idempotent. But with the high throughput of some systems I was working on, even redis with its atomic nature of put and get and handy dandy sets and lists, would have race conditions.
Now I hardly touch redis because my skills in making sure things get processed just once have improved greatly. And I can just use PostgreSQL and my code.
Being Idempotent from the beginning is now the center of my life.
Similarly, some part of the system remains imperative. The network card at least, will always resend a packet. The goal is to pass around idempotent messages except for the very leaves.
For instance, instead of an endpoint 'email(customer, data)' you might have 'email_or_report_on_send_status(customer, data)' and the later endpoint would check the cache for (customer, data) and merely report the previous results if it found them.
I agree though, this stuff used to keep me up at night and eventually I've grown more natural about not mutating things unless I mean it. (This phrase sounds like a comic-book villain.)
If you’re making or using an API where repeating would be bad, consider using idempotency tokens for those too. I believe Stripe supports them. The basic idea is the same: if you pass a token into them, they will guarantee that in a certain time frame, no other requests with that ID can be duplicated. This is useful when the network flakes during the response. Is it safe to retry?
Things get trickier when you combine network and database consistency measures; that’s when you get into locks and multi stage commit and etc. and it helps to know your database’s consistency model, since it’s often not as solid as you think! (In the past, even PostgreSQL had issues with providing serializable isolation.)
That's handled by the extra condition in step 1 of the altered rules:
> which haven't been charged the fee this month.
A couple of key aspects are:
1. The payments API must be idempotent for the Cron to be idempotent. Otherwise they can't safely retry timeout failures.
2. The operations <charge-the-customer> and <mark-the-customer-as-charged> have to be (to loosely use the term) atomic.
For #1, a payment processor should explicitly state idempotency semantics. Look at Stripe [1] for example. They require you to pass Idempotency Key header in charge request. Clients must carefully choose this key to fit the needs for their business flow and use case. For the author since they charge monthly the idempotency key must be such that it should be constant for a given customer and given month. So something like "customerid-YYYY-MM" makes sense. Now no matter which day of the month cron runs the customer is guaranteed to charge at most once.
As for #2 most of data stores now a days support atomic operations so it should be reasonably straight forward. Get-and-set in Redis, start/end transaction with "for update" in MySQL etc.,
Note that #2 is optional if you don't mind spurious retries to Stripe.
Stepping back from the details it can be noticed that for a system/operation to be idempotent all its downstream dependencies should support idempotent operations. In this example, if Stripe weren't idempotent then you have to resort to manual fixes if charge operation times out.
Another observation is, if a downstream system can assure idempotency then don't bother making your operation idempotent. Failed payment? Retry. No need to store if a payment was attempted or not. Make sure the idempotency-key generation logic is robust.
PS: Source - I've been in payments industry for a while so spend a lot of time thinking about and implementing idempotency. The latest one being when my team worked on payments API at Uber. We included a idempotency in the glossary [2] section ;-)
[1] https://stripe.com/docs/api/idempotent_requests [2] https://developer.uber.com/docs/payments/glossary
Payments at Google defined some specs for payment companies to build to, which allows easy onboarding as a form of payment on google's platform. We have a page talking about idempotency and expected behavior.
https://developers.google.com/standard-payments/reference/id...
I don't think it's that different from other payment companies, like adyen or worldpay, but we attempted to explain our view of idempotency on a single page.
Edit: Even mentioning the rule to someone who asked can bring out the haters.
This is actually a lot better than previous companies. My employee contract when I was at Cisco, basically said I couldn't talk about anything Cisco related on the internet without sign off from upper management or PR or something.
If I have an idempotent method like "CreateCustomerRecord", this can cause a lot of pain for audit features and other aspects of the domain model if it is internally making determinations about whether to actually create or silently skip creation. For me, I would much rather that the method throw an exception if there is a duplicate business key than have it silently complete without taking any actual action. Exceptions indicating attempts at invalid state transitions can be extremely valuable if you have the discipline to create & use them properly.
Generally, seeking idempotence in otherwise mutable methods is a band-aid for when you have broken immutability rules and allowed things to leak out of the sacred garden of unit-tested state machines and other provably-correct code items.
If you should only conditionally execute some method, perhaps the solution is to investigate the caller(s) of the method, rather than attempt to infer the intent of all possible callers within the method itself.
Idempotency is great for "debouncing" requests. If you want to tell difference between identical requests that are different transactions, add a unique transaction id of some kind.
(In this particular example, I agree that an idempotency token is probably the way to go, as otherwise callers would need to somehow distinguish between their callers trying to create a duplicate account, versus something going wrong with them resulting in duplicate requests. I just have too often seen developers conflate "idempotency" with "use an idempotency token)
I can honestly say that Eric is 100% right with his approach. It always leads to less headaches, more flexibility (oh trust me, someone is always gonna have a "but... there's like a special thing that I sometimes have to do" and it breaks some assumptions.
In any case... yeah... let's just say any time you have to be worried "did we already schedule this", really think "can this never care if it was or not? Should be always safe to schedule it again"
"Idempotence is the property of a software that when run 1 or more times, it only has the effect of being run once."
And the example is, instead of a chron job just running a process once a month or on some other schedule, it runs more frequently but checks if the change has already been made.
(From the latin Idem which means "same" and potence is of course power/potent, so it has the same power/effect however many times you run it)
Nilpotence (x^2 = 0) is also very helpful some times: it's a process which is self-reversing. Like the discrete Fourier transform (if you set up the constants properly).
Squaring a upper triangular matrix with 0 on the diagonal is nilpotent. Derivatiting a polynomial of degree N is nilpotent after N iteration.
Also you may be confusing it with x^n = 1 (which I'm not sure how to name, 'root of unity' perhaps). This would be the case for the Fourier transform (with n=4).
If x^2 = 0 then applying the Fourier transform twice would null your function, which isn't the case.
f(x) = f(f(x)) = f f x = f^2 x = f x
And leaving out the x, because it is just a placeholder anyways:
f f = f^2 = f
And of course this means that
f^n = f because f^n-1 f = f^n-1 by induction.
Particularly in this case O^2 = O makes more sense to me than f(f(x)) = f(x). And I just naturally think about A B - B A = 0 rather than f(g(x)) - g(f(x)) = 0 for commutativity.
In mathematics, a self-reversing function is called an involution, and it's f^2 (or f(f) ) = Id, the identity function.
Nilpotence is very different. It says that if you apply your function a certain number of times, you end up with zero no matter what the input is. For example, projection on x axis + 90 ° rotation of a vector is nilpotent.
The value in programs being a "free rerun" was that every so often the program would barf on a bad bit of data in a record.
The programming environemnt was interpreted BASIC, so if an error occurred the program would print a message on the console and drop to an interactive prompt.
The operators running the batch schedule would see this and call the programmer on call for that night. You'd log in (over dial up at this time) and attach to the process, look at the error, figure out what went wrong, either correct the data or (more likely) skip the record and deal with it the next day. It was more important to have the programs finish on time; individual issues could be dealt with later.
Often you could just start up the program from where it left off, but if things were more screwed up it was important to be able to re-run it without any negative consequence.
Edit: this was ~30 years ago, so my point is that it's not any kind of new idea or something that wasn't recognized long ago.
And a lot of "secret" scripts that automated a big part of our job.
As a property, I think it's even nicer if a script can literally fully run twice and for the outcome to be the same if it only ran once (so skipping the 'did I run before?' check).
Even though this check is useful in general, if you can define your data in such a way if it did somehow run, that this is not destructive / creates incorrect data, it makes the system more robust.
Of course this is not always possible though. For example, if the process results in an email being sent, you need an explicit check to not do that twice.
So the goal of your functional core is to fully construct the email, and return it to the caller, who then has the choice to send the email, print it, write it to disk, etc.
The program you deploy looks like this
EmailSender().send_email(construct_email(args))
You can test by implementing a "safe" EmailSender interface, so that you're executing the same code that's in prod.
In general, if a job/function is mutating state deep in the syntax tree (i.e. sending emails in the middle of a batch job), I personally see that as a violation of the Single Responsibility Principle.
> We need to charge dormant customers a monthly fee so that we don't have to keep their money on our accounting books forever.
That's absolutely and objectively not the reason this happens. Keeping their money "on [your] books" has zero cost associated with it outside of the escheatment[0] process, which can be completely automated. I worked at a company that had to do a decent amount of escheatment for almost every US state, Canada, and Mexico, and for a total company size of under ten people the entire organization probably spent 45 minutes working on it in any given month. It's a trivial amount of time.
This happens to that the company can turn it's customers' money into its money. I don't know enough about to escheatment process to know for sure but I'd imagine regular debits like this would also get around the need to return the money to the estate after a given period of time. So someone puts $500 in their account, gets hit by a bus the next day, and the company gets to take a little bit every month/quarter/year until it's all theirs.
It's a pretty scummy thing to do and I'm surprised the author just takes it at face value that they "have to" do it.
[0] https://en.wikipedia.org/wiki/Escheat#Transfer_agents_and_es...
> 2. Charge each of these accounts a fee
> 3. Setup a cron job to run this every hour
Note that if this job ever runs successfully, but takes more than an hour, you will double-count. Can easily happen if the box running these crons is overloaded. One fix is to automatically halt the job after 55 minutes, another would be to have the middle step be impotent, for each user you're doing the process on, ensure (ideally in a threadsafe manner) that they need the operation to be done still.
If some rogue deploy script creates fifty of the cron job, and weirder things happen every day, one of those could be checking the stale state during the transaction of another one. A classic data race.
Eliminating this problem is left as an exercise for the interested reader...
1. Balance is not a simple variable, but the sum of all credits and debits to an account
2. A fee is a charge record in your database
3. This fee has a database constraint that you can have only one record per month
Now you can run the script that charges dormant fees as often as you want.Practically speaking, this narrows your "transaction window" significantly - instead of:
1. Begin transaction
2. Check to see if work has already been done
3. Do work
4. Persist work
5. Commit
With a potentially long transaction spanning from 1-5, you do this: 1. Do work
2. Persist work to key / table with uniqueness constraint
3. On conflict, do nothing (looks like you already did the work before)
Of course if "Do work" is very expensive, you can bring back in "Check to see if work has already been done" as an optimization, but for many simple CRUD examples, it's actually /cheaper/ to learn that the work has already been done via the conflict check failing than via an explicit pre-flight check.> 2. Charge each of these accounts a fee.
There's still a race condition between 1. and 2. if you don't atomically check the account hasn't been charged yet as you're charging it.
With long running processes like these it's also pretty much guaranteed to happen if you accidentally run two instances of this process at the same time.
As an engineer, the more exposure and experience you get, the more insight you have about the ways things can fail. Identifying the ways something can fail is the really important step here. You can’t know what failsafe to implement if you don’t actually know how something can fail. But once you do know how something can fail, implementing a proper solution is easy a lot of the time, even if you are not explicitly aware of the concept of idempotence, for instance. Only in some tricky Byzantine edge cases do you need very specific, well-established, track-proven algorithmic solutions.
If you can build a system with ACID 2.0 life gets really easy. You can reason about your system without worrying about ordering, time, 'exactly once' semantics, etc.
Idempotency is usually one of the simplest pieces to implement, and you definitely get a ton of benefit right off the bat - it's worth designing systems from scratch with it in mind.
Using mutation iterator, last update time and knowing the rate of change over time you can back apply any any calculations that have been missed regardless of how long it’s been since your cron job ran.
I setup a server to run the job on startup and 1 time a day at midnight. This also allows multiple different charges and interests to be applied to the same account. I also set it up to apply this process every time an account was accessed to make sure the accounts were up to date.
The idempotence is nice and I use it as much as possible, but, his case is very simplistic. Imagine you have a complex data migration process, which crashes in the middle - how you safely continue from where you crashed or safely roll back?
"If table doesn't have column X, add it. If row Y doesn't exist in table Z, add it." etc.
Article made me think it's actually applicable to other management as well
UPSERTs can be idempotent as well. "If this doesn't exist, create it, and if it does, update it to match this state", implies that running it twice will leave no unintended side effects.
Now that I think about it, all the situations where I'm fine using POST are also idempotent, such as changing an object's metadata (technically PUT or PATCH), or sending a batch of HTTP requests to an endpoint bundled as a single request.
A more interesting function is PUT (idempotent) vs POST (not idempotent)
Also, in web front end programming, there are lots of cases where X needs to cause Y to happen, but Z also needs to cause Y to happen. It's much much simpler if Y is idempotent, rather than X checking if Z has already happened etc.
My reasoning is that if a table represents the current state of the world, then any state changes should be made with UPSERT in order to bring the table up to date with the world. If a table represents the history of changes, then that history should be appended to with INSERT, but not modified.
I'm sure that there are other cases that I'm not currently considering, but I'm also a newbie at SQL and would love to be told of them.
One impact of this is that is forces the changes to be atomic. If an object-oriented interface updates a database whenever properties are changed, then it results in many single-column changes. In the example below, if there is a network failure between the two commands, then changing the name from "Jane Doe" to "John Smith" could result in an unintended intermediate state of "John Doe".
customer_obj.set_forename("John")
customer_obj.set_surname("Smith")
(Again, I am by no means an expert, and in part posted this so that I can be corrected if/where I am wrong.)I'm skeptical this is desirable even just from a security perspective.
I think you're considering the role of the database in a narrower context than they're used. Certainly the databases I work with, it's a minority of updates (actually a very small minority) that are performed as a result of a user inputting something into a form in the sense you're talking (i.e. updating an entity). Think of how many updates are of the form update this order to change status to shipped or product stock available to X. You wouldn't want to be obligated to pass the entire row of something to an application so it could pass back a single changed column.
This is not to even get into the scenario where an update is applied to a view (which might be a limited result from one or more tables). In that sort of scenario it's not even clear what an UPSERT would even do.
#1 : collect the state the "user" (might be you) desires the system to be in.
#2 : collect the state the system is actually in.
#3 : compare #1 to #2 and see if you're already there and if you are, you're done.
#4 : if they differ, then do the work once to fix them.
#5 : bonus: go collect the system state again and send it out to report on what changed
Interesting I know floating point arithmetic is not associative, but I think it is communitative.
This was his only leading principle. Result: absolute chaos - the code aspired to be idempotent, but due to idempotency he avoided thinking problems through and just created a mess of individual functions - each being idempotent, aside from the unavoidable bugs - which didn't form a coherent flow at all.
We did a major refactoring, threw out about all that code, rewrote everything in a logical manner. Now everything is still idempotent, but comprehensible.
TLDR: idempotency is the same snakeoil as the majority of guiding principles: alone, it doesn't help at all. There are lots of other factors to consider, which make the developer/architect role demanding (and fun).
Craftmanship at least, a sense for architecture (better) or understanding the whole picture of the requirements as a team of developers (best) is still required.
A bit of pragmatism goes a long way, like Python's odd
x = (some_tuple,)
. . .syntax amidst its generally clean approach.
Inflexibility itself is the bugaboo.