anyone got some insight? is this just a toy?
anyone got some insight? is this just a toy?
If you're asking whether the catch clause in a Temporal Workflow saga is guaranteed to execute, the answer is yes. The way it's able to guarantee this is by persisting each step the code takes and recovering program state if a process crashes, server loses power, etc. For an explanation of how this works, see: https://temporal.io/blog/building-reliable-distributed-syste...
Generally, that's a good way of thinking about it. The one additional bit of nuance is it's like a "safe" stack unwind while other processes could be still modifying databases at the same time, so it's not a complete "rollback" of the whole world if that makes sense.
Also related: Signals are events that you can send to Workflows and between Workflows, and they’re always delivered in the order they’re received.
More generally, for a handy reference of Distributed Systems patterns, check out https://microservices.io/patterns/data/saga.html (though I personally find his diagrams a bit...overwhelming) and the MSN writeups: https://learn.microsoft.com/en-us/azure/architecture/pattern...
You don't need to explicitly interact with an event queue because it's a higher level abstraction that sits on a queue. In fact, that’s the big value add.
ACID doesn't help you once you're trying to coordinate actions across systems, like the example in the article.
The event queue is an even bigger value add:
* It's the audit log
* A human reading the event queue has final say over the 'true state' of the system (insofar as such a thing exists in a distributed system)
There are several Saga-related application-specific patterns called countermeasures, which help to somewhat mitigate this problem.
Also unlike RDBMS with their ACID transactions, Saga design forces you to understand your business requirements better. Which steps are more likely to fail? Which steps are pivotal (i.e. points of no return)? Which steps are riskier or more valuable for business, etc.