HNHacker News
TopNewBestAskShowJobs

gvdongen

4 karma · joined December 11, 2023

submissionscomments
gvdongen··on Temporal raises $550M at a $12.55B valuation
Yes it does: https://docs.restate.dev/services/deploy/cloudflare-workers
gvdongen··on Temporal raises $550M at a $12.55B valuation
If a handler starts a sleep/human approval/RPC call, or so, then this timer/promise is persisted in Restate's journal. Restate does the waiting. The handler process itself can suspend (e.g. on a serverless function), and the bidirectional connection is closed. Once the timer fires/approval comes in, Restate re-invokes the service with the journal of previously completed steps, and the service can replay to the exact point in the code where it suspended and continue from there. Restate is like a DB for journals, so you can sleep for as long as needed, also months.

So you have fast persistence of events while a handler can make progress, and suspensions while waiting.

gvdongen··on Temporal raises $550M at a $12.55B valuation
Hi! I work for Restate

A few key differences. Restate has a more flexible programming model. You don't write workflows with activitities, but just durable processes/handlers. Durable steps execute inline and get persisted over an open streaming connection in Restate (low latency, lower overhead per durable step, sharing resources like sandboxes) instead of working with a pull-model where each activity executes remotely on a worker. Restate has a lean deployment model with a single binary that can be deployed multiple times to have a highly-available cluster (potentially spread across multiple regions). It is used for large-scale production clusters, and so lightweight here does not mean less reliable than Temporal.

You can do the same things with Temporal like sleep for months etc. You can learn more here: https://restate.dev/vs/temporal

gvdongen··on Agent checkpointing is far from production-grade resiliency
As agents run longer and spend more money, many agent frameworks are adding resiliency features like checkpoint recovery and pause-resume approvals.

But to get your agent to production, checkpointing is not enough. There is quite a big gap left for you to handle: failure detection, automatic retries, high availability, scale-out, idempotency, concurrency, session coordination, versioning, ...

I wrote a blog post on what’s left to solve, and how to solve it.

TL;DR Instead of tying resiliency together with your agent framework, agents should be built on top of a highly-available orchestration layer that owns the end-to-end execution, guarantees it completes, and handles all of the points above.

Optionally, agent frameworks can be used on top of this to help abstracting away the agent loop.

Is this also how you see it and productionize your agents?

gvdongen··on Updating AI Agents safely in production
AI agents often run for minutes, hours, or even days. This makes deploying new versions tricky: what happens if an agent starts on one version of your code and resumes later after you’ve deployed another?

If you swap out code from under an agent, execution can break. Or worse, a changed description or tool implementation can cause your agent to silently misinterpret its own history and produce inconsistent results.

This blog post introduces how to solve this via immutable deployments and pinned executions: - Each deployment lives at a unique endpoint and represents an immutable, versioned snapshot of your code, prompts, tools, and schemas - Every execution is pinned to its deployment; retries, resumptions, and callbacks always return to the same version - New requests route to latest

The blog post shows how Restate implements this in practice. Versioning becomes an infrastructure property you don’t need to think about, rather than something you solve in code.

gvdongen··on Every System is a Log: Avoiding coordination in distributed applications
We've used purely excalidraw. Nice to hear you like them!
gvdongen··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Here is a fan-out example for async tasks: https://docs.restate.dev/use-cases/async-tasks#parallelizing... First, a number of tasks are scheduled, and then their results are collected (fan-in). This probably comes closest to what you are looking for. Each of those tasks gets executed durably, and their execution tracked by Restate.
gvdongen··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Here is another example in the examples repo which does compensation. There is also a Java one https://github.com/restatedev/examples/blob/main/basics/basi...