Distributed transactions? Two stage commits? Just do it and rely on the fact you have 99.9% uptime and it's _probably_ not going to fail?
Anyone else dealt with this headache?
Distributed transactions? Two stage commits? Just do it and rely on the fact you have 99.9% uptime and it's _probably_ not going to fail?
Anyone else dealt with this headache?
And, since it's microservices, it's near impossible to refactor it, while it could have been a simple thing to reorganize some code in a monolith (where it is also a good idea to make sure that DB transactions don't span wildly different parts of the source code, but the refactor to make that happen is easier).
This is the big downside of microservices -- not the difficulty of doing atomic operations, but the difficulty of changing the architecture once you realize you drew the wrong service boundaries.
Microservices is great as long as you choose the perfect service boundaries when you start. To me, that's like saying you always write bug-free code the first time -- it's not doable in practice for large complex projects -- hence I'm not a fan of microservices...
An example being something like an online gun store, you have a perfect service that handles orders. It's completely isolated and works fine. But now, 2 years later some local government has asked you "whenever someone buys a gun, you need to call into our webservice the moment the order is placed so we know a gun was sold, and you need to successfully do it, or you can't sell the item"
Now you've got a situation where you need an atomic operation, place the order and call the web service, or don't do it at all. You could say just place the order, do the web service call asynchronously and then delete the order afterwards. But you might not have that choice depending on what the regulations say. You can't do it before you place the order because what if payment fails?
The order service should not have any idea about blocking orders in specific scenarios. And now the architecture has broken down. Do you add this to the order service and break it's single responsibility? Will this be a bigger problem in the future and do you need to completely rearchitect your solution?
Even with a monolith, what if you have a power-off at the wrong moment?
What you are describing here is to me pretty much the job description of a backend programmer to me -- think through and prepare for what happens if power disappears between code line N and code line N+1 in all situations.
In your specific example one would probably use a reserve/capture flow with the payment services provider; first get a reservation for the amount, then do the external webservice call, then finally to a capture call.
In our code we pretty much always write "I am about to call external webservice" to our database in one DB transaction (as an event), then call the external webservice, and finally if we get a response, write "I am done calling external webservice" as an event. And then there's a background worker that sits and monitors for cases of "about-to-call events without matching completed-events within 5 minutes", and does according required actions to clean up.
If a monolith "solves" this problem then I would say the monolith is buggy. A monolith should also be able to always have a sudden power-off without misbehaving.
A power-off between line N and N+1 in a monolith is pretty much the same as a call between two microservices failing at the wrong moment. Not a qualitative difference only a quantitive one (in that power-off MAY be more rare than network errors).
Where the difference is is in the things that an ACID database allows you to commit atomically (changes to your internal data either all happening or none happening).
I think that's one of the things people rarely think of when moving to microservices. Just how much effort needs to be made to rectify errors.
You can always guarantee atomicity. You will just have to implement it yourself (what is not easy, but always possible, unless there are conflicting requisites of performance and network distribution).
And yes, the cleanup jobs are part of how you implement it. But you shouldn't be "trying to rectify the problems", you should be rectifying the problems, with certainty.
Create the order, and set the status to pending. Keep checking the web service until it allows the transaction. Set the status to authorized and set the status to payment. Keep trying the payment until it succeeds, officially set the order and set status to ordered.
I really find it hard to believe that the regulations won't allow you to check the web service for authorization before creating the order. If that's really the case then create the order and check, and if it doesn't work, then cancel the status of the order and retry. It's only a few rows in the database. If this happens often then show the data to your local politician or what not and tell them they need to add more flexibility to the regulation.
a) redraw the transaction boundaries, aka. avoid it, or
b) don't do it (see below),
c) idempotency -- so you can just retry everything until you succeed.
You can do distributed transactions, but it's a huge pain operationally.
(b) is not always as far fetched as you might think. The vast majority of Real World (e.g. inventory, order fulfilment generally, etc.) systems are actually much more like Eventual Consistency in practice, so you don't gain anything by being transactional when your Source of Truth is a real world thing. Inventory is a great example. It doesn't matter if your DB says that an item is in stock if there simply isn't an item on the shelf when it comes time to fulfill an order for that item. Tracking such things is always imperfect and there already other systems overlaid which can handle the failure case, e.g. refund the order, place on backorder, have a CSR find a replacement item, etc. etc. (Of course you want reasonable approximate values for analytics, logistics, etc.)
Fortunately there are not that many things in the world that need to be 100% atomic so you can get away with a lot.
For your own microservices you generally have at least the option of "fixing" the problem properly even if it's at great expense.
But then you hit external systems and the problem resurfaces.
You can go crazy thinking about this stuff, at a certain point most business logic starts to look like connectors keeping different weird databases in sync, often poorly.
Pure crud api? Oh that's a database where client is responsible for orchestration (create folder, upload document...) and some operations might be atomic but there are no transactions for you. Also, the atomicity of any particular operation is not actually guaranteed so it could change next week.
Sending an email or an SMS? You're committing to a far away "database" but actually you never know if the commit was successful or not.
Payments are a weird one. You can do this perfectly good distributed transaction and then it fails months after it succeeded!
Travel booking? runs away screaming so many databases.
etc.
Sagas are about recognizing the individual changes that are necessary, and dealing with the success or failure of them at a higher level. This is complicated though as the developer and the business now need to have a specific conversation around what happens if A succeeds and B fails? Does A need to get "rolled back"? Does B need to be retried? Does C need to wait until B succeeds before proceeding? That all brings in a level of complexity and the only answer is to manage and use the appropriate patterns and tools to do so. Wanting to go back to a land where you can just wrap it all in a transaction so that you get one nice boolean indicating success at the end is quite frankly just naïve. The real world doesn't work like that.
They're both gonna do their own database functions.