A few key differences. Restate has a more flexible programming model. You don't write workflows with activitities, but just durable processes/handlers. Durable steps execute inline and get persisted over an open streaming connection in Restate (low latency, lower overhead per durable step, sharing resources like sandboxes) instead of working with a pull-model where each activity executes remotely on a worker. Restate has a lean deployment model with a single binary that can be deployed multiple times to have a highly-available cluster (potentially spread across multiple regions). It is used for large-scale production clusters, and so lightweight here does not mean less reliable than Temporal.
You can do the same things with Temporal like sleep for months etc. You can learn more here: https://restate.dev/vs/temporal
Naively without digging into the code I would look at a “streaming over an open connection” as likely to strictly more brittle.
So you have fast persistence of events while a handler can make progress, and suspensions while waiting.
While the workflow is actively doing work (like calling other services, accessing state or interacting with external services, for example) the server is connected to the workflow deployment via a low-latency bidirectional streaming connection to receive and acknowledge progress that the workflow makes. That way the workflow can finish as fast as possible.
A nice side effect of this model is that you can co-locate your workflow with expensive resources, such as a sandbox, which should be used by all durable actions that the workflow executes. The reason this works is because Restate can inline durable steps (what would be modeled in Temporal as activities, I believe).
While the durable function does fast work and adds steps, it pushes it through a stream. When a wait point comes, it closes and replays on resumption (typical durable execution style).
That gives you the best of both worlds: same long-running workflows with long sleets and suspensions, but also ability to add steps with few ms overhead only.
Temporal has 3 types of activities:
* local: executed in the same process as the orchestrator code. Many local activities can be executed locally before their results are sent to a backend server in a single RPC call.
* task queue based: executed by a pool of worker processes that poll from the queue. This is the most flexible model as it supports flow control, priorities, fair queueing out of the box.
* eager dispatch: task queue based but executed locally if possible as a performance optimization.
Temporal also supports stand alone activities that are invoked without a workflow and dispatched through a task queue.
Restate only supports local activities (using Temporal terminology).
I wouldn't call it a "more flexible programming model".
Restate made several decisions I consider questionable for the system's availability and stability, like pushing work to handlers instead of dispatching it through a queue. Any design decision has tradeoffs. It would be nice if you mentioned these trade-offs in your posts instead of making claims that sound like pure marketing.
(1) You can model the equivalent of local activities and activities that run on other workers in Restate.
A local activity is a step in the workflow function. An activity supposed to run on a different worker is a function called by the workflow function. Since these calls are just Restate events, exactly-one, suspendable, this gives you a full-fledged workflow/remote-activity pattern. Including concurrency, separate retry policies, etc.
(2) Restate steps commit individually, unlike local activities.
Imagine a two-step workflow, where you want one step durable before starting the second. Account withdrawal before deposit. Restate steps allow you to do that, each step is durable committed before the next step runs.
Per Temporal's own docs, Temporal Local Activity results become durable only when the enclosing Workflow Task completes. That's different than Restate, which can durably commit every individual ctx.run before proceeding to the next step.
Making actual durable commits fast, so you can have sequences of fast durable steps building on each other is super valuable. If an agent can commit the guardrail evaluation in low milliseconds before it starts the tool call, that's great, do it. If committing this involves dispatching another workflow or activity task, the consideration is harder.
When we see someone migrate a Temporal workflow, they often end up using many more durable steps in Restate than they used activities before.
(3) Why do we consider it flexible?
(a) Virtual Objects: Keep state around across the workflow, without doing tricks like "keep the workflow running, signal only, continue_as_new" after a while. Virtual Objects are a natural way to model concurrent stateful entities.
(b) non-workflow communication patterns: We have seen users build lot's of different patterns. It can get as crazy as graphs of functions/objects sending each other durable messages. All end-to-end idempotent (or exactly once, for Virtual Object state). That is outside the hierarchical workflow/subworkflow/activity abstraction.
(4) availability and stability
I don't know where the perception with stability comes from, Restate pushes some pretty high volumes for customers, like 100k+ actions/sec.
The push model in Restate is internally dispatched through queues as well. The application just don't see it as task queues. Limits are implemented through virtual-queues in Restate 1.7.
-----
Of course the systems make different trade-offs. The Restate design (vqueues, push model) is took us longer to build than a task queue model would have, because it puts more work onto the dispatcher that is otherwise handled just implicitly by the worker pools.
But once it is there, it is sooo nice, in how easy it integrates into infra, and how it can handle flow-control with high hierarchical limits in ways I genuinely haven't seen in any other system achieve (see https://restate.dev/blog/announcing-restate-1-7 )
Stay humble / nice ego dude.
Necessary reference: https://share.google/7uReiRkd6mY3RcCQQ
I certainly don't think I'm superior to them in any way and have great respect for the very capable team they are. We know each other personally. I even presented at the Flink conference while at Uber.
My message is that I'd prefer a more technical discussion of the merits of our products rather than simplified marketing attacks.
Sure, Temporal is more mature and battle-tested, but Restate is quite nice and I honestly prefer it in most cases.
EDIT:
I started drafting this reply before seeing someone from the Restate team chimed in.
I'm sure there is a place for restate.dev, but it isn't in the same place as Temporal.
Like Temporal, Restate supports workflows that sleep for months at a time and let's them reliably finish.
It actually does many more useful things. For example, you can have services that remember state across invocations and there is no need for continue-as-new.
Since Restate allows inlining durable steps into your workflow, it is very easy to co-locate your workflows with expensive resources such as sandboxes or other per-node resources. And the nice thing is that each durable step is really cheap and adds only minimal overhead.
Restate's spiritual father is Stateful Functions (https://nightlies.apache.org/flink/flink-statefun-docs-maste...) a library for event driven applications built on top of Apache Flink. If you want to learn more about why we started building Restate, I recommend this excellent blog post https://restate.dev/blog/why-we-built-restate.
I understand that Restate and Temporal look similar from a superficial perspective but Restate is not only a durable execution engine but a durable runtime that also provides consistent state and reliable communication. These are the building blocks to build reliable agents and applications w/o having to fit them into a workflow-activity like model. If you want to learn more about how Temporal differs from Restate, check out https://restate.dev/vs/temporal.