HNHacker News
TopNewBestAskShowJobs

sewen

188 karma · joined January 13, 2015

submissionscomments
sewen··on Temporal raises $550M at a $12.55B valuation
Hmmm, that feels not very nice, tbh. I have big respect for what Maxim and the team built, and both systems are essentially open and free (publish code and are free to use).
sewen··on Temporal raises $550M at a $12.55B valuation
Sorry, but this is incorrect (Restate founder here)

(1) You can model the equivalent of local activities and activities that run on other workers in Restate.

A local activity is a step in the workflow function. An activity supposed to run on a different worker is a function called by the workflow function. Since these calls are just Restate events, exactly-one, suspendable, this gives you a full-fledged workflow/remote-activity pattern. Including concurrency, separate retry policies, etc.

(2) Restate steps commit individually, unlike local activities.

Imagine a two-step workflow, where you want one step durable before starting the second. Account withdrawal before deposit. Restate steps allow you to do that, each step is durable committed before the next step runs.

Per Temporal's own docs, Temporal Local Activity results become durable only when the enclosing Workflow Task completes. That's different than Restate, which can durably commit every individual ctx.run before proceeding to the next step.

Making actual durable commits fast, so you can have sequences of fast durable steps building on each other is super valuable. If an agent can commit the guardrail evaluation in low milliseconds before it starts the tool call, that's great, do it. If committing this involves dispatching another workflow or activity task, the consideration is harder.

When we see someone migrate a Temporal workflow, they often end up using many more durable steps in Restate than they used activities before.

(3) Why do we consider it flexible?

(a) Virtual Objects: Keep state around across the workflow, without doing tricks like "keep the workflow running, signal only, continue_as_new" after a while. Virtual Objects are a natural way to model concurrent stateful entities.

(b) non-workflow communication patterns: We have seen users build lot's of different patterns. It can get as crazy as graphs of functions/objects sending each other durable messages. All end-to-end idempotent (or exactly once, for Virtual Object state). That is outside the hierarchical workflow/subworkflow/activity abstraction.

(4) availability and stability

I don't know where the perception with stability comes from, Restate pushes some pretty high volumes for customers, like 100k+ actions/sec.

The push model in Restate is internally dispatched through queues as well. The application just don't see it as task queues. Limits are implemented through virtual-queues in Restate 1.7.

-----

Of course the systems make different trade-offs. The Restate design (vqueues, push model) is took us longer to build than a task queue model would have, because it puts more work onto the dispatcher that is otherwise handled just implicitly by the worker pools.

But once it is there, it is sooo nice, in how easy it integrates into infra, and how it can handle flow-control with high hierarchical limits in ways I genuinely haven't seen in any other system achieve (see https://restate.dev/blog/announcing-restate-1-7 )

sewen··on Temporal raises $550M at a $12.55B valuation
That's fair, but they are also a great fit for Restate. For example, Replit migrated their whole coding agent from Temporal over to Restate, a pretty big setup.
sewen··on Temporal raises $550M at a $12.55B valuation
I guess the confusion is that it doesn't have to be one single persistent stream.

While the durable function does fast work and adds steps, it pushes it through a stream. When a wait point comes, it closes and replays on resumption (typical durable execution style).

That gives you the best of both worlds: same long-running workflows with long sleets and suspensions, but also ability to add steps with few ms overhead only.

sewen··on [dead]
We tried building a scalable and resilient cloud coding agent with @restatedev for workflows, @modal for sandboxes, @vercel for compute, and GPT-5 / Claude as the LLM. Think a mini version of cursor background agents or Lovable, with a focus on scalability, resilience, orchestration.

It is a fun exercise, there are many interesting patterns and micro problems to solve: * Durable steps (no retry spaghetti) * Managing session across workflows (remember conversations) * Interrupting an ongoing coding task to add new context * Robust life cycles for resources (sandboxes) * Scalable serverless deployments * Tracing / replay / metrics

Sharing our learnings here for anyone who builds agents at scale, beyond "hello world"

sewen··on Building a modern durable execution engine from first principles
I just realized I missed an important part: The primary durability for the bulk of the state comes from S3 (or similar object store). The periodic snapshots give you like an automatic frequent backup mechanism for free, which in itself is a nice property to have.
sewen··on Building a modern durable execution engine from first principles
Indeed, the persistence layer is sensitive, and we do take this pretty serious.

All data is persisted via RocksDB. Not only the materialized state of invocations and journals, but even the log itself uses RocksDB as the storage layer for sequence of events. We do that to benefit from the insane testing and hardening that Meta has done (they run millions of instances). We are currently even trying to understand which operations and code paths Meta uses most, to adopt the code to use those, to get the best-tested paths possible.

The more sensitive part would be the consensus log, which only comes into play if you run in distributed deployments. In a way, that puts us into a similar boat as companies like Neon: having reliably single-node storage engine, but having to build the replication and failover around that. But in that is also the value-add over most databases.

We do actually use Jepsen internally for a lot of testing.

(Side note: Jepsen might be one of the most valuable things that this industry has - the value it adds cannot be overstated)

sewen··on Building a modern durable execution engine from first principles
afaik, with Temporal you deploy workers. When a workflow calls an activity, the activity gets added to a queue, and the workers pull activities from queues.

In Restate, there are no workers like that. The durable functions (which contain the equivalent of the activity logic) get deployed on FaaS or like a containerized RPC service. The Restate broker calls the function/service with the argument and some attached context (journal, state, ...).

You can think of it a bit like Kafka vs. EventBridge. The former needs long lived clients that poll for events, the latter pushes events to subscribers/listeners.

This "push" (Restate broker calls the service) means there doesn't have to be a long running process waiting for work (by polling a queue).

I think the difference also naturally from the programming abstraction: In Temporal, it is workflows that create activities, in Restate it stateful durable functions (bundled into services).

sewen··on Building a modern durable execution engine from first principles
Here is a comparison to Temporal, maybe that helps with a comparison to those systems as well? https://news.ycombinator.com/item?id=43511814
sewen··on Building a modern durable execution engine from first principles
There are a few dimensions where this is different.

(1) The design is a fully self-contained stack, event-driven, with its own replicated log and embedded storage engine.

That lets it ship as a single binary that you can use without dependency (on your laptop or the cloud). It is really easy to run.

It also scales out by starting more nodes. Every layer scales hand-in hand, from log to processors. (you should give it an object store to offload data, when running distributed)

The goal is a really simple and lightweight way to run yourself, while incrementally scaling to very large setups when necessary. I think that is non-trivial to do with most other systems.

(2) Restate pushes events, compared to Temporal pulling activities. This is to some extent a matter of taste, though the push model has a way to work very naturally with serverless functions (lambda, CF workers, fly.io, ...).

(3) Restate models services and stateful functions, not workflows. This means you can model logic that keeps state for longer than what would be the scope of a workflow (you have like a K/V store transactionally integrated with durable executions). It also supports RPC and messaging between functions (exactly-once integrated with the durable execution).

(4) The event-driven runtime, together with the push model, gets fairly good latencies (low overhead of durable execution).

sewen··on Building a modern durable execution engine from first principles
The way we think about durable execution is that it is not just for long-running code, where you may want to suspend and later resume. In those cases, low-latency implementations would not matter, agreed.

But durable execution is immensely helpful for anything that has multiple steps that build on each other. Anytime your service interacts with multiple APIs, updates some state, keeps locks, or queues events. Payment processing, inventory, order processing, ledgers, token issuing, etc. Almost all backend logic that changes state ultimately benefits from a durable execution foundation. The database stores the business data, but there is so much implicit orchestration/coordination-related state - having a durable execution foundation makes all of this so much easier to reason about.

The question is then: Can we make the overhead low enough and the system lightweight enough such that it becomes attractive to use it for all those cases? That's what we are trying to build here.

sewen··on Building a modern durable execution engine from first principles
All of the Restate co-founders com from various stages of Apache Flink.

Restate is in many ways a mirror image to Flink. Both are event-streaming architectures, but otherwise make a lot of contrary design choices.

(This is not really helpful to understand what Restate does for you, but it is an interesting tid bit about the design.)

       Flink     |   Restate
  -------------------------------
                 |
    analytics    |  transactions
                 |
  coarse-grained |  fine-grained
    snapshots    | quorum replication
                 |
   throughput-   |  latency-sensitive
    optimized    |  
                 |
  app and Flink- |  disaggregated code
  share process  |   and framework
                 |
      Java       |      Rust
the list goes on...
sewen··on Building a modern durable execution engine from first principles
Thank you for the kind words!

The storage engine is pretty tightly integrated with the log, but the programming model allows you to attach quasi arbitrary state to keys.

So see whether this fits your use case, would be great to better understand the data and structure you are working with. Do you have a link where we could look at this?

sewen··on The Anatomy of a Durable Execution Stack from First Principles
The post discusses the design considerations when building a durable execution runtime from the ground up.

The goal is a highly-available, transactional, scalable, and low latency runtime in a self-contained binary that scales from laptop to complex distributed deployment.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Yes, there is one, have a look at https://restate.dev/cloud/
sewen··on Every System is a Log: Avoiding coordination in distributed applications
This is certainly building on principles and ideas from a long history of computer science research.

And yes, there are moment where you go "oh, we implicitly gave up xyz (i.e., causal order across steps) when we started adopting architecture pqr (microservices). But here is a thought on how to bring that back without breaking the benefits of pqr".

If you want, you can think of this as one of these cases. I would argue that there is tremendous practical value in that (I found that to be the case throughout my career).

And technology advances in zig zag lines. You add capability x but lose y on the way and later someone finds a way to have x and y together. That's progress.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Great question:

The Virtual Objects in Restate are much like actors. They are somewhat inspired by Orleans [1], and you could call them virtual stateful actors. They blend with the durable execution for processing messages with multiple durable steps.

Regarding temporal, check also this question: https://news.ycombinator.com/item?id=42815318

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Temporal is related, but I would say it is a subset of this.

If you only consider appending results of steps of a handler, then you have something like Temporal.

This here uses the log also for RPC between services, for state that outlives an individual handler execution (state that outlives a workflow, in Temporal's terms).

sewen··on Every System is a Log: Avoiding coordination in distributed applications
You can catch these errors and handle them in a common try/catch manner, and because the results of `ctx.run` are recorded in the log, this is deterministic and reliable
sewen··on Every System is a Log: Avoiding coordination in distributed applications
I can see where some of that could be written more clearly. To elaborate:

- We mean using one log across different concerns like state a, communication with b, lock c. Often that is in the scope of a single entity (payment, user, session, etc.) and thus the scope for the one log is still small, and it reduces coordination headache for coordinating between the systems. You would have a lot of independent logs still, for separate payments.

- It does _not_ mean that one should share the same log (and partition) for all the entities in your app, like necessarily funneling all users, payments, etc. through the same log. What would be needed if you try and do some multi-key-distributed transaction processing. That goes actually beyond the proposal here, and has some benefits of its own, but have a hard time scaling.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Some clarification on what "one log" means here:

- It means using one log across different concerns like state a, communication with b, lock c. Often that is in the scope of a single entity (payment, user, session, etc.) and thus the scope for the one log is still small. You would have a lot of independent logs still, for separate payments.

- It does _not_ mean that one should share the same log (and partition) for all the entities in your app, like necessarily funneling all users, payments, etc. through the same log. That goes actually beyond the proposal here - has some benefits of its own, but have a hard time scaling.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Yes, we are assuming a log that picks linearizability at the cost of availability under partitions. Like most logs do, including Kafka, Pulsar, RedPanda, etc.

The application state is defined by the log here, and the log drives retries/recovery, so it doesn't much matter if the process that executes the app code splits off. The log would hydrate another one.

Also the one log is at the granularity of a single key or handler execution. More of a logical log, than a physical log or even partition.

In Restate, we implement a logical log-per-key, backed by a partitioned physical log.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Yes, exactly right. One log per logical entity, here "payment ID".

The way our open source project implements that is with a partitioned log and indexes at key-granularity, so it is like virtually a log per key.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
There is nothing to coordinate for the application, because, yes, the log coordinates everything. But not globally, on the level of a single event handler execution, or a single key.

That has been proven to scale well - the way we implement that in Restate is classical shared nothing physical partitioning, with indexing on a key granularity.

So nothing like a shared mutex unless you want to access the same key, which otherwise your database synchronizes, if you want any reasonable level of consistency.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
[2] https://martin.kleppmann.com/2015/11/05/database-inside-out-...
sewen··on Every System is a Log: Avoiding coordination in distributed applications
That blog post is a great read as well. Truely, the log abstraction [1] and "Turning the DB inside out" [2] have been hugely influential.

In a way this article here suggests to extend that

(1) from a log that represents data (upserts, cdc, etc.) to a log of coordination commands (update this, acquire that log, journal that steo)

(2) have a way to link the events related to a broader operation (handler execution) together

(3) make the log aware of handler execution (better yet, put it in charge), so you can automatically fence outdated executions

[1] https://engineering.linkedin.com/distributed-systems/log-wha...

sewen··on Every System is a Log: Avoiding coordination in distributed applications
I assume CSP is communicating sequential processes?

Interesting analogy - in a way it is doing something CSP-like in a distributed app/service architecture with the all the different processes and components that are there. The shared log (or a partition of that) being a way to establish a sequential order.

sewen··on Every System is a Log: Avoiding coordination in distributed applications
That gist is correct - I would add that the log needs a few specific properties and conceptually be the shared log for state, communication, execution scheduling.

The next step is the, how do you make this usable in practice...

sewen··on Every System is a Log: Avoiding coordination in distributed applications
Haha, no, but maybe all the AI-generated contents out there is starting to train me to write in a similar style...
sewen··on Every System is a Log: Avoiding coordination in distributed applications
Never encountered it before, but it looks cool.

I think they are trying to solve a related problem. "We can consolidate the work by making a generic log that has networking and syncing built-in. This can be used by developers to make automatically-decentralized apps without writing a single line of networking code."

At a first glance, I would say that Gossiplog is a bit more low level, targeting developers of databases and queues, to save them from re-building a log every time. But then there are elements of sharing the log between components. Worth a deeper look, but seems a bit lower level abstraction.

Page 1 of 2Next →