Let's say you're making a plane booking service (kiwi.com clone).
You charge the customer, then book the flight. But if booking the flight fails, you must refund the customer.
try {
pay();
} catch {
try {
refund();
} catch {
// "must refund the customer" implies we can't reach here
}
}Step1
Step2
Step1Undo
then this has a 1% chance of needing manual repair (it's okay if step1 fails, but if step1 succeeds and step2 fails, we need to repair):
do Step1
do Step2
and this has a .01% chance (we only repair if Step2 and Step1Undo fails, 1% * 1%):
do Step1
try {
do Step2
} catch { do Step1Undo
}In case Step1's service doesn't expose an API to poll its status, then the only recourse is to execute it again (with the same input key, assuming it's idempotent ;)
1. You book a flight. You successfully reserve a seat.
2. You book a car. You successfully reserve a sedan.
3. You try to book a hotel room. The room that you wanted was booked while you were booking your flight, and there aren't any more available.
You obviously don't want the car or flight anymore, and you want to cancel them without a human having to manually fix it.
The answer is, you model those as well and work out what to do. But it's more messy than you might think if you just model the first-order failure paths.
1. A misunderstanding of the business rules. In the flight example, you thought that were flights were cancellable, but actually the airline only offers nonrefundable seats.
2. System type errors, e.g. network outages.
If you get a type 1 failure, that's an error that gets ingested in your error monitoring service, and is a bug that needs to be fixed. If you get a type 2 failure, idempotent cancellation (which is necessary for this work) will eventually get you to your desired state. Either way, you shouldn't need to model deeper into the state graph.
That would have been a good article. The saga pattern could have just been a footnote to it.
Instead of untangling the mess, just cut the gordian knot and throw a nice error of what failed and what was aborted.
So instead of having a complex logic. Have a simple lambda function that talks to a queue. That's it. It accepts an undo command. You read a command you stuff it in the queue done. No DB, No servers. If you were running this yourself. You will have a simple API (distributed) that does the same to a distributed queue/cache. Done.
Your complex job can now pick up the undo commands from the queue and execute with logic to retry if for some reason it fails.
If you have if/else in your code, you don't think "If one path fail, so can the other path, so I never handle the other path".
Regardless of whatever happens, a failed transaction state should always be possible without affecting data integrity.
It's just an act of trying to automate the resolution of error scenarios to reduce human effort.
Billing is one, for example.