Temporal .NET – Deterministic Workflow Authoring in .NET
temporal.io
temporal.io
So it's great for any code that you want to ensure reliably runs, but having methods that can't fail also opens up new possibilities, like you can:
- Write a method that implements a subscription, charging a card and sleeping for 30 days in a loop. The `await Workflow.DelayAsync(TimeSpan.FromDays(30))` is transparently translated into a persisted timer that will continue executing the method when it goes off, and in the meantime doesn't consume resources beyond the timer record in the database.
- Store data in variables instead of a database, because you can trust that the variables will be accurate for the duration of the method execution, and execution spans server restarts!
- Write methods that last indefinitely and model an entity, like a customer, that maintains their loyalty program points in an instance variable. (Workflows can receive RPCs called Signals and Queries for sending data to the method ("User just made a purchase for $30, so please add 300 loyalty points") and getting data out from the method ("What's the user's points total?").
- Write a saga that maintains consistency across services / data stores without manually setting up choreography or orchestration, with a simple try/catch statement. (A workflow method is like automatic orchestration.)
There are multiple strategies to update live workflows, see https://community.temporal.io/t/workflow-versioning-strategi.... Most often for small changes, people would use the `Workflow.Patched` API.
That doesn’t take into account the time it’s already been waiting. For when you want to do something on a schedule and be able to edit the schedule in future, there’s a Schedules feature that allows you to periodically start a workflow, like a cron but more flexible. [2] In this case, the workflow code would be simpler, like just ChargeCustomer() and SendEmailNotification().
- Creates and updates events/messages for you
- Maintains consistency between the events and timers and the state of the processing flow. More info on why this is important for correctness/resilience: https://youtu.be/t524U9CixZ0
The difference between workflow code and using an event bus is that with the former, the above is done automatically for you, and in the latter, it's done manually, which can be a lot of code to write and get right, and is harder to track/visualize what happened in production and debug. It would also take a lot of events to get an equivalent degree of reliability—the message processor would need to do a single step and then write the next event to the bus. So a 10-line workflow-as-code function would translate to 10 different events in the bus route.
Also, the event bus route doesn't have the new possibilities I listed in the parent comment.
Sure you can claim this is a fundamentally lower level mechanism, but really it's the same thing.
It is probably important, like all enterprise software quagmires, to consider how they are sold.
Your typical IT manager with low code skills has all his documented high level processes his group manages. Hey look, Visio diagrams with point to point flows of boxes with arrows, and ... maybe ... some interactions of state with databases and/or systems domains.
They know just enough that the actual nuts and bolts is buried in all this ... code. Actual unintelligible code to them, even if they have some inkling of coding.
To them, it's simple flow, and his PHBs above him can understand what he does with the simple flows.
Then some workflow vendor walks in the door, pops up some visual editor, and wows him with basically "you don't need all that code, you put it into the workflow tool and it will just work, and BOOM your coding environment IS your simple visio document".
WOW SIGN ME UP TAKE $$$$$$$! Then comes the pilot flows.
Error handling and retries?
Distributed State and even worse, Transactional update of Distributed state?
Load distribution?
Branching, looping?
Systems integration?
It goes back to the theory of computation. State machines have limitations in processing. The stack machine addresses some of that, but it eventually runs out, requiring ... the turing machine.
As it turns out, almost all enterprise data flows or processes require turing machines. That's why they are coded at some level by turning complete languages.
Superficially at a high level, you start to see a basic state machine model on top of that ... but it is an illusion.
You move the turing machine into the workflow engine (and the workflow engine IS a turing machine ... they all have them: state, looping, branching) and the "simple point to point" flow becomes spaghetti ... tool-locked in spaghetti, with fixed limits on ability to do things.
The current evolution to workflows is the "directed acyclic graph" workflow engine. This has been an improvement, mostly by constraining the actual use of workflow engines to task organizations that they can do, and trying to keep people from going "full Turing" in the workflow engine.
It still can loop ... most do it by recursive calls to subflows ... gets pretty spaghetti. And you still have the fundamental issue that all PHBs will want in the workflow. On error, retry, or have a recovery flow, or that type. Still a huge amount of complexity to properly get the workflow working.
And yet the visual editing workflow tool can have enormous value. Enormous. The Visual nature of the flow, ability to visually diagnose suspended / failed executions. And workflow are everywhere: batch processes, code builds, deployments, automated maintenance, backups / restores, etc.
And I haven't even gotten into the mess of automated rules-engine-based stuff.
The only value structuring low level code along the lines of what "enterprise workflow" has evolved into after decades (useful, but not a holy grail) is if it gives you a fundamentally better way to visualize the execution of the code, which can happen under constrained use of workflow engines.
UML was a massive disaster, another tangential relation to what appears to be being done here. There your "workflows" or code diagrams were code generated to code.
Alas, the final problem of workflow engines is their balkanization. XML standardizations (BPEL) failed miserably for all the usual corporate product standardizations (functioned as lockin for the existing players, lowest common denominator abilities, ugly, XKCD protocol+1).
If only... if only there was a good designed representation scheme and a wide variety of good open source visualization and execution engines. But there aren't.
I think what is discussed here is a step towards a potential solution: it comes from the IDE tooling, something that a workflow always was (in the vein of the now-defunct CASE/Computer Aided Software Engineering days). A standard tooling that coders demand and IDEs provide as a minimum barrier. But IDEs are single machine things, and workflows are distributed entities ... sigh, nevermind that thought.
Ok, maybe we just need a good visualization tool first that is more universal. Don't care about the creation, just something that can "plug in" and represent non-workflow system interactions AS workflows. "Enterprise execution visualization". A REALLY good system for that has never existed IMO, and is universally needed.
> The Visual nature of the flow, ability to visually diagnose suspended / failed executions.
Temporal has a web UI in which you can see which executions are failing, and see on which step they're failing:
I think they also have C# SDK which I can't vouch for because I haven't used it.
There’s a lot to learn with it. I’ve seen ramp-up take a few months per engineer, though we’re also making it a little harder on ourselves by self-hosting, being the early adopters on the Python SDK (Go, Java, and TypeScript are the most mature I think), and dealing with a mix of Python async and multiprocessing (a bunch of CPU bound activities in the mix). The docs are solid, and the team is responsive to community users.
Glad to see that Temporal follows a similar approach and gets you the same benefits. For every coder out there currently using AWS SWF: if your day work involves more than just handholding a handful workflows but building those, take a look at Temporal. You'll never look back.
(To be fair I am still grumpy that they still separate "deciders" and activitities, but I can see the benefits of that.)
If you want to use a GUI to design workflows: equally useful, but probably with a different target audience.
The Typescript SDK is amazing. It strikes a nice balance, being straightforward to use without most of the common pitfalls related to non-determinism, thanks to the Temporal SDK using Node's VM module to execute code with patched sources of non-determinism, like `Date` and `math.random`.
The only caveat is running it (Temporal Server). After mulling it over, for this project I sticked with running everything in a single VPS. If for some reason I get more users one day, maybe consider a Kubernetes managed service. Temporal Cloud is also an option, just not for me at the moment (region constraints, plus it's a pet project so no money)
But the local developer experience is actually amazing. Temporal is a joy all around, to be honest.
What would be the downside of writing it such as:
`ExecuteActivityAsync<OneClickBuyWorkflow>(e => e.DoPurchaseAsync, etc...)`?
'We solve this problem by allowing users to create instances of the class/interface without invoking anything on it. For classes this is done via FormatterServices.GetUninitializedObject and for interfaces this is done via Castle Dynamic Proxy. This lets us "reference" (hence the name Ref) methods on the objects without actually instantiating them with side effects. Method calls should never be made on these objects (and most wouldn't work anyways).'
Which sounds like a lot of heavy lifting. It seems something like
public class WorkflowBuilder<T> where T : class
{
// Somehow workflow gets the real instance
private T Instance;
public async Task<TResult> ExecuteActivityAsync<TResult>(Func<T, Task<TResult>> func) => await func(Instance);
public async Task<TResult> ExecuteActivityAsync<TResult>(Func<Task<TResult>> func) => await func();
public async Task ExecuteActivityAsync(Func<Task> func) => await func();
public async Task ExecuteActivityAsync(Func<T, Task> func) => await func(Instance);
}
public record Purchase(string ItemID, string UserID);
public class PurchaseActivities
{
public static WorkflowBuilder<PurchaseActivities> OneClickBuyWorkflow => new WorkflowBuilder<PurchaseActivities>();
public async Task DoPurchaseAsync(Purchase purchase)
{
await OneClickBuyWorkflow.ExecuteActivityAsync(e => e.DoPurchaseAsync(purchase));
}
public static async Task DoPurchaseAsyncStatic(Purchase purchase)
{
await OneClickBuyWorkflow.ExecuteActivityAsync(() => DoPurchaseAsyncStatic(purchase));
}
}
Would pretty much achieve the same thing.How is `e` created for the `e => e.DoPurchaseAsync(purchase)` lambda? You're going to have to do that lifting anyways to create an instance for `e` that isn't really a usable instance. Unless you use source generators which we plan on doing.
I think what you have there is a lot more heavy lifting. Also note that workflows and activities are unrelated to each other. Workflow can invoke any activities. The code you have is a bit confusing because `PurchaseActivities` should be completely unrelated to workflows.
But this should never happen, these is no reason to create an unusable instance. The real instance should be resolved in the same way however the current workflow resolves it.
You can jump through a bunch of hoops like requiring interfaces which is what some frameworks do. But in our case, we just decided to make it easy to reference the method without invoking it or its instance.
You can parse an expression to serialize it and run it on a different server etc
See e.g. https://github.com/6bee/Remote.Linq
Of course the "Ref" pattern is user-choice/suggested-pattern, it's not a requirement in any of our calls that just take simple delegates however you can create those delegates. So I may be able to work it in there.
E.g. https://docs.hangfire.io/en/latest/background-methods/passin...
BackgroundJob.Enqueue<EmailSender>(x => x.Send(13, "Hello!"));One big difference is those three are Python only, and Temporal is Go, Java, JS/TypeScript, Python, PHP, and .NET. Usually another difference between most workflow tooling and Temporal (granted I have not checked with these) is the care they take in handling errors, retry, replay, cancel, etc so that'd be worth checking as well.
It's for writing code that has steps that involve humans. E.g.: sending a mail notification for someone to approve a document, that kind of thing.
These are inherently slow and asynchronous, because humans operate on timescales of hours or days, not milliseconds.
Temporal, while it uses the "workflow" terminology, is a new type of thing. At a basic level, it's "do you want your backend code to run reliably?" If yes, and you're okay with the latency hit of each step getting persisted for you, then the answer is "use Temporal to write your backend code." It's a new programming model that lets you develop at a higher level of abstraction, where you don't have to be concerned about faults in the hardware or network or downstream services/3rd party APIs being temporarily down. Where you no longer have to code retries, timeouts, use task queues, or use a message bus to communicate between services. And oftentimes don't even need to use a database.