Saga Pattern Made Easy
temporal.io
temporal.io
What did the API of the meeting service look like specifically?
Billing is one, for example.
1. You book a flight. You successfully reserve a seat.
2. You book a car. You successfully reserve a sedan.
3. You try to book a hotel room. The room that you wanted was booked while you were booking your flight, and there aren't any more available.
You obviously don't want the car or flight anymore, and you want to cancel them without a human having to manually fix it.
The answer is, you model those as well and work out what to do. But it's more messy than you might think if you just model the first-order failure paths.
Instead of untangling the mess, just cut the gordian knot and throw a nice error of what failed and what was aborted.
1. A misunderstanding of the business rules. In the flight example, you thought that were flights were cancellable, but actually the airline only offers nonrefundable seats.
2. System type errors, e.g. network outages.
If you get a type 1 failure, that's an error that gets ingested in your error monitoring service, and is a bug that needs to be fixed. If you get a type 2 failure, idempotent cancellation (which is necessary for this work) will eventually get you to your desired state. Either way, you shouldn't need to model deeper into the state graph.
That would have been a good article. The saga pattern could have just been a footnote to it.
So instead of having a complex logic. Have a simple lambda function that talks to a queue. That's it. It accepts an undo command. You read a command you stuff it in the queue done. No DB, No servers. If you were running this yourself. You will have a simple API (distributed) that does the same to a distributed queue/cache. Done.
Your complex job can now pick up the undo commands from the queue and execute with logic to retry if for some reason it fails.
Let's say you're making a plane booking service (kiwi.com clone).
You charge the customer, then book the flight. But if booking the flight fails, you must refund the customer.
try {
pay();
} catch {
try {
refund();
} catch {
// "must refund the customer" implies we can't reach here
}
}Regardless of whatever happens, a failed transaction state should always be possible without affecting data integrity.
Step1
Step2
Step1Undo
then this has a 1% chance of needing manual repair (it's okay if step1 fails, but if step1 succeeds and step2 fails, we need to repair):
do Step1
do Step2
and this has a .01% chance (we only repair if Step2 and Step1Undo fails, 1% * 1%):
do Step1
try {
do Step2
} catch { do Step1Undo
}In case Step1's service doesn't expose an API to poll its status, then the only recourse is to execute it again (with the same input key, assuming it's idempotent ;)
It's just an act of trying to automate the resolution of error scenarios to reduce human effort.
If you have if/else in your code, you don't think "If one path fail, so can the other path, so I never handle the other path".
Suppose temporal is correctly configured and operated, and gives you reliable and robust workflow execution. Your custom workflow executes code that attempts to perform operations on your other services over the network, via those services' APIs. Temporal itself doesn't know or model anything about your API calls, or your other services. Temporal uses the abstraction of an "activity" to wrap operations with side-effects, such as API calls to your services that workflow might need to make. Temporal can guarantee that it will keep re-trying to execute an activity until it succeeds, but this means that the code in your activity may get executed multiple times.
If you want each state-changing operation in your workflow to execute exactly-once, so your overall system has exactly-only execution semantics, it is your responsibility to ensure that the APIs exposed by your other services and used in your workflow are idempotent, and that they are called correctly from the workflow, passing an idempotency key that does not vary when a temporal "action" wrapping an API call is retried.
For another interesting (temporal-agnostic) discussion of how to build a system achieving similar 'exactly once' behaviour, there's a good infoq podcast interview with Jason Maude: https://www.infoq.com/podcasts/cloud-based-banking-startup-j...
The situation is similar to the whole "database inside out" hype that drives application programmers to re-implement proper DBMS.
I am pretty sure the IT hype wheel will turn around and someone will sell 2PC as a new shiny thing.
However, if all of the data you need to update is in a single database that supports atomic commits, I'd go with that over sagas.
Scalability depends on lock granularity, what's more...
> Sagas [...] can have high throughput
There is no real difference as sagas in practice implement locking in disguise - take the scenario of flight/hotel/car booking:
Once you book a hotel - this particular resource (a room at a particular time) is effectively locked. Cancelling the booking (because other participants failed in saga) is effectively releasing the lock. The room (resource) is locked for the duration of the whole process anyway (as no other customer can book it during this time).
The downside of sagas is that a programmer is forced to explicitly handle all failure scenarios - which costs development time, is error-prone etc.
anyone got some insight? is this just a toy?
If you're asking whether the catch clause in a Temporal Workflow saga is guaranteed to execute, the answer is yes. The way it's able to guarantee this is by persisting each step the code takes and recovering program state if a process crashes, server loses power, etc. For an explanation of how this works, see: https://temporal.io/blog/building-reliable-distributed-syste...
Generally, that's a good way of thinking about it. The one additional bit of nuance is it's like a "safe" stack unwind while other processes could be still modifying databases at the same time, so it's not a complete "rollback" of the whole world if that makes sense.
Also related: Signals are events that you can send to Workflows and between Workflows, and they’re always delivered in the order they’re received.
More generally, for a handy reference of Distributed Systems patterns, check out https://microservices.io/patterns/data/saga.html (though I personally find his diagrams a bit...overwhelming) and the MSN writeups: https://learn.microsoft.com/en-us/azure/architecture/pattern...
You don't need to explicitly interact with an event queue because it's a higher level abstraction that sits on a queue. In fact, that’s the big value add.
ACID doesn't help you once you're trying to coordinate actions across systems, like the example in the article.
The event queue is an even bigger value add:
* It's the audit log
* A human reading the event queue has final say over the 'true state' of the system (insofar as such a thing exists in a distributed system)
There are several Saga-related application-specific patterns called countermeasures, which help to somewhat mitigate this problem.
Also unlike RDBMS with their ACID transactions, Saga design forces you to understand your business requirements better. Which steps are more likely to fail? Which steps are pivotal (i.e. points of no return)? Which steps are riskier or more valuable for business, etc.