Distributed transactions are not microservices
talentica.com
talentica.com
But as with anything, the question is "is this complexity worth the price you pay for it?" For me, given the advantages that microservices offer in many contexts, I find that using the Saga Pattern to maintain consistent state is totally worth it. But that won't be true for everybody in every situation.
Or you do a two-phase commit: A, B, and C tentatively succeed, but then one of the commit calls fails, and now?
It seems like inconsistencies are inevitable no matter what you do.
At some level, barring guaranteed message delivery (which is effectively non-existent in any distributed system) you always reach a level where you can't guarantee consistency. It's the Byzantine General's Problem, basically.
https://en.wikipedia.org/wiki/Byzantine_fault
But based on empirical evidence, you can work out that a certain measure of effort dedicated to fault tolerance will yield correct results in X% of cases, and you can tune the value of X based on how much time/energy/money/effort you're willing to expend... up to a point.
One gotcha that is not covered by Sagas (I could be wrong) is when one or many of the network paths involved in the distributed tx become unreachable (network partition event) and you have no idea of the state of that part of the tx. Do you re-try that part and risk sending the same instruction twice (ok in some cases but not all) vs risk of having sent no instruction? If I had to implement a distributed tx I would first verify my mental model using TLA+ and use a (persistent) transactional messaging system with at-least-once delivery as the backbone, and make other accommodations for such scenarios.
If you can make your compensating action idempotent, then yes, you can just keep retrying it. If it can't be made so for whatever reason, then a failure at that point demands manual intervention.
Accounting has been doing that for centuries already, so it's not new by any means. It's also not free, it imposes severe restrictions on your system's architecture and the kinds of problems it can solve.
I think the premise of using a monolithic service (your API "Controller") to handle transactions that it simply shouldn't be concerned with is the main problem here. i.e the problem is that the transaction should not be distributed, not that it is handled through multiple services.
I'm further confused at the reconciliation that seems to occur later in your example. Why is it bypassing the services and writing directly to the service DB itself?
You can tar an awful lot of microservices with that brush.
If someone at VP-level is making low-level tech decisions, GTFO. If your non-technical executive management even wants to know what the low-level tech driving their business is, GTFO. If your manager, Director, VP, execs, etc will not listen to honest, calm "we really don't need ________ because {5 rational, evidence-backed reasons}," GTFO.
If there really was one answer, and the answer was just as simple as "get rid of the whole management team, and you'll have a much better product at the end" then I have to imagine companies would have started doing this already. My experience being on both sides of this coin in my career is that: it's just not that easy.
If the CEO is the lead developer and the head of IT and the CFO what signs all them checks, that's obviously different, and in that case everyone should probably understand everything. Otherwise, ask yourself how deep management gets into the minutiae of washing the toilets.
Either:
Set the inventory amount in your e-commerce system to be less than the actual inventory (which is rarely accurate anyway). This is your safety stock and depends on how fast moving the item is and if it's a close-out that you are really trying to sell to zero. Then just handle exceptions at allocation-time when you're able to commit stock.
Or:
Avoid a two-phase commit problem by allocating stock at add-to-cart time with a ticketing system that allows a hold to be placed with a timeout. This is a more customer friendly approach that handles stampedes better, such as caused by marketing emails.
Either way, inventory management is like banking, aiming to be eventually consistent is a lot more realistic than being always consistent.
Yeah I love those basic DB examples of transactions and why it's important using bank accounts when in real life banks are all eventually consistent because they had to solve the problem before distributed transactions were available.
https://github.com/seata/seata
https://wecode.wepay.com/posts/waltz-a-distributed-write-ahe...
At the end of the day I don't think anyone should be coding distributed transactions into their app's. If you need use a solution that abstracts them for you. Eventually we will get one into Vitess ( https://vitess.io ) when we have figured out something general purpose enough to work for lots of workloads
Starbucks does not use two-phase commit
https://www.enterpriseintegrationpatterns.com/ramblings/18_s...
If you can't rely on an actors intent to such a degree, your actions need to also be less costly to be efficient.
Correlated Ids and idempotent end points with retry is pretty much a two phase commit. Something eventually checks and deals with exceptions.
Accounting for loss and developing ways to operate with it make a process viable.
What’s the arbitrator pattern? Google failed me.