Temporal: open-source microservices orchestration platform
temporal.io
temporal.io
I'm excited for Maxim, Samar and team :)
Let's say I have some workflow which executes over a period of days or weeks, and I want to deploy a new version of it. How does that work? What happens to workflow instances that are currently executing?
After writing the above I clicked back to the site to see if this was addressed, and found this https://docs.temporal.io/docs/java-versioning and... surely you're not serious?
The standard approach of versioning the whole workflow is OK for introducing new features but doesn't really support bug fixing. For example, a workflow is expected to run for three months. A bug is found at the end of this workflow definition. The bug is fixed and deployed as a new version, but all workflows that started up to the fix and used previous versions will keep failing for the next three months. So there is a need to patch workflow without changing its version.
The approach Temporal takes is that every part of the code is versioned independently. This allows deploying changes at any time, even for libraries shared by multiple workflows, and doesn't require running multiple worker versions.
Put another way, it's a great big leak in the abstraction.
It doesn't end up in spaghetti code as old branches are proactively removed after they are not needed anymore. Temporal provides APIs that allow to count number of workflows using each version.
I don't think it is a leaky abstraction as it is indeed a part of the business logic of deciding how old version of the state should become a new one. I don't think there is a generic solution that allows implicitly migrate old state to a new state on any long running computation.
A workflow is not just externally stored data with a well-defined schema. It is the computation's whole state, including threads blocked on API calls and local variables on the stack. So it is not possible to define translations between different versions the way Edit Lenses do.
When I read the article "Why the Serverless Revolution Has Stalled" [1] on here a couple weeks ago, my reaction was that the reason is that serverless doesn't solve the business logic issue. Serverless removes operational details of servers, but often exacerbates process fragmentation. And the need to have "serverful" operational expertise is replaced by the need for serverless expertise. At least right now, this new expertise is far from trivial.
I haven't yet used Temporal, but I've spent a lot of time evaluating it, and its predecessor, Cadence. The idea is to model long-running business logic essentially as procedures, in ordinary code. These are called Workflows, and must be free of external effects. External effects are carried out by Activities, and the scheduling and tracking of results are handled by the Temporal runtime. The upshot is that you get to write workflows as though they have no time constraints. It's like async programming, but liberated from the confines of a OS process or machine.
If it turns out to be a useful home for business logic (I understand that it has at Uber), I think the next frontier is integrating it within UI frameworks. I'm imagining the next Rails being something like Next.js + Temporal. I still have a bunch of questions in my mind, like how to decide which data lives in Temporal vs. a OLTP database. Someone with more experience using Temporal probably could better answer this.
One of the reasons I'm particularly interested in this topic is because my company, Better.com, uses our own homebuilt workflow engine to model the days-long, multi-user business project of mortgage origination. In our case, our workflow engine is actually built as part of a full-stack framework that goes well beyond Temporal in scope, but it's not built as a generic platform, and we look at Temporal for inspiration on where things might be going.
[1] https://www.infoq.com/articles/serverless-stalled/
[2] I discuss this a bit here https://www.themuse.com/advice/engineering-manager-better-al...
It also has a great visual interface using BPMN for understanding where your processes are in a flow. I don't see anything like that in Temporal.
https://camunda.com/why-camunda/
Incidentally, an iteration of Camunda called Zeebe leverages event-sourcing to work more efficiently at scale.
I've used it for microservices orchestration previously and it works great.
I'm excited for Zeebe too but it's a bit immature atm.
Temporal represents all the business logic in one place in the programming language of your choice. All parameter passing is strongly typed.
I see people advocating visual representation for workflows. But if it is so great for programming in general, why they don't advocate the same for systems programming, for example? Linux kernel in JSON/XML anyone?
BPMN vs code is something we have considered. Our workflow engine works more like state machines as code, which is kind of the worst of both worlds. The business processes get fragmented into logic for each state, and mixed with imperative effects, so its neither easy to see the process nor maintain the code.
The nice thing about BPMN is that it's probably a bit more learnable by the subject matter experts who ultimately determine what the correct business process should be. However, having seen some pretty monstrous process flowcharts, my guess is that it probably doesn't scale past a certain level of process complexity. I suspect that workflow-as-code is ultimately the more scalable option.
[1] https://www.infoq.com/articles/events-workflow-automation/
waitForApproval(purchase);
sleep(Duration.ofDays(30));
sendEmail(email);I totally agree. As a consultant working with multiple companies, I see a slowly increasing number of Serverless applications. The code always looked bloated, with no chance of ever being migrated off the original cloud platform due to all the proprietary APIs in use. There was also no way to run any of them locally without extensive configuration. Serverless has a future, but this is not it.
Talking of "Serverless Revolution", I think we are going to see more of such abstractions as things evolve. Abstractions and tools built upon serverless functions that are going to cater more closely to problems being solved rather than worrying about managing new-found complexity of the functions themselves.
BPMN spec has been around for more than a decade, and there are many implementations. What does Temporal bring to the table that is different/better then existing BPMN engines?
* Strongly typed (for languages like Java or Go)
* Uses standard error handling of the language of choice. For example, Java SDK throws exceptions, and Go one returns errors.
* Use standard tools for development like IDEs, Debuggers, Linters, Unit Testing frameworks, etc.
* Allows using the same language for both activities and workflows.
* Allows programs of practically unlimited complexity through standard programming language techniques like OO, functions, etc.
* Easily supports handling of asynchronous events
* Supports updating definitions of already running workflows
* Allows reuse of standard libraries. For example, if workflow needs to keep a state in a priority queue, it can use an existing one.
My personal opinion that BPMN tries to cater to two distinct groups of people. Non-technical domain experts and software developers. And it is essentially a compromise that cannot serve both of them well. Non-technical domain experts still cannot implement production-ready workflows, and developers have to use limited complicated UI/XML-based technology that is inferior to the programming languages and environments they are used to. The Temporal approach gives developers the best tool for their job and lets them decide how to interact with non-technical users. In some domains, it is possible to create DSL for non-technical users to use. Temporal is an excellent technology for implementing such domain-specific languages. BPMN's problem is that it is not domain specific and tries to serve as a general-purpose language without being one.
Do you plan to maintain both projects, or plan to deprecate Cadence in favor of Temporal?
We spent almost year working on various improvements before declaring the first production release. The most important technical difference is that Temporal uses gRPC/Protobuf when Cadence is TChannel/Thrift. One implication is that TChannel did not support any security and Temporal supports mTLS. We are currently working on a comprehensive security story.
Here is a stackoverflow post that has more details about this: https://stackoverflow.com/questions/61157400/temporal-workfl....
I ask because internally we just deployed our first Cadence cluster a few weeks ago -- it'd be good to know what to expect and what the founders of the project suggest to do now that Temporal is officially released to production?
The long answer is that we see Cadence as a great part of our lineage. We felt that our vision for the open-source of Temporal was much different than what Uber sees for the project.
Now that we have a production grade release, we are slowly phasing out Cadence support. We will still provide free support to any Cadence users who need help migrating to Temporal.
I can’t speak to your situation without knowing more. I will say if you would ever like a hosted version of this tech, Temporal is the way to go. I am happy to answer any specific questions too!
ryland@temporal.io
Disclaimer: head of product at Temporal. Temporal is not container orchestration and is not an infrastructure management tool. In most cases users run Temporal on top of Kubernetes.
Temporal provides a distributed experience which is decoupled from the reliability of any specific piece of hardware. We provide a programming model for writing distributed applications without needing to code around all points of failure. We still are working on what to call the tech exactly, it’s not something that is widely known by any name today. It’s sort of like virtualized distributed computing.
That being said, Temporal backend consists of a few stateless and horizontally scalable services (matching service, frontend service etc). Because these roles experience load differently it often makes sense to scale them separately. Due to this design, users often find it convenient to use an orchestration solution such as Kubernetes, ECS etc. HashiCorp themselves run our technology using Nomad to directly answer your question.
The only thing we are strongly opinionated about is that you run the underlying database in a production-grade manner. Throwing a MySQL container into a helm chart isn't going to cut it for serious usage.