Pi Durable
earendil.com
earendil.com
The main reasons are:
1) they are "durable", i.e. easier to make long-running in an unattended way, and easier to implement recovery, monitoring, etc
2) separating the harness from the compute brings safety and scaling benefits
3) easier to make multi-player.
I like your structured concurrency approach with tasks, which is similar to how I do it in Lightspeed too.
Also, the durable state implementation as documents is elegant! Question, though: why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?
We tried so many things. At one point it pulls in so much more complexity. At one point we had half of automerge's proxy system in there. In the end we felt like this is a reasonable line to draw, but we will see!
The good thing is that this more low level API can be easily papered over with a nice sugary thing.
Pi-durable makes pi a good acquisition target for Cloudflare - nothing like durable objects (with containers no less) really exists in other clouds. Wonder what the_mitsuhiko thinks about this.
> acquisition target for Cloudflare
Once it gets acquired, it won't be any better than all the others.
Pi is alright but it's only virtue so far is being the Neovim of harnesses, minimal yet (incredibly) extensible. That being said it's still early days and unclear how the whole landscape regarding agentic stuff will play out at various different levels. For example I prefer stuff like Claude Code and Pi but I've seen a friend use Kiro at work with some crazy workflows all basically structured around Markdown files, spec writing, ingesting tickets from Jira and then validating/testing the code written automagically.
So, it'll be interesting to see if these durable agent setups need more of a complete product approach, or if people want to compose them as libraries.
The durability has been a lifesaver when I'm dogfooding and the TUI crashes or the session spawns 5 ambitious subagents and OOMs my laptop. Quite a few times I've migrated a session to Cloud's engines and kept going from there. Effectively the entire state of multiple sandboxes is put through storage as humble OTLP data and revitalized on whatever hardware you bring it back on, like thawing Walt Disney (and giving him a stimpack I guess).
My end goal is to have long-running sessions that hold curated context via tool state so it's durable to compaction, and to be able to keep a multi-agent month-long workstream going from my phone by messaging its leader through Cloud. It's been a passion project for over a year and it's taken a lot of world-building along the way (new Dagger core APIs, and Dagger 1.0 being the top priority independent from all this agent stuff). I'm finally at the point where I'm using it productively but I don't have a quick-start yet (edit: here's a quick and dirty one - https://gist.github.com/vito/bed465ed43a09a433b85a360311c7d3...)
Find me in Dagger.io's Discord if you're interested :)
I'm still using Concourse every day... for fun (and GPU mutex) — I guess I'll switch to Dagger? If it's light enough to run agents, I'm very curious.
Definitely joining the Discord server.
In Pinthe coding agent, each transcript entry is parented to another entry. That was actually exceptionally dumb.
If you do /tree in pi, pi needs to flatten that tree into linear, nested conversations.
In Pi Durable, we corrected this mistake. A conversation is a chronological, immutable list of entries. A conversation can be parented to an entry in another conversation, and thus inherits that parent's older conversation entries starting from that entry.
So, exactly the same functionality, just less dumb.
Like if I have a web-app running on the runner and the agent is navigating the web UI and then the runner (or the agent) crashes. When the agent is recreated back from the checkpoints (or a new runner is launched), it will think it has already navigated to page N, but in reality the browser on the runner might be on page 0.
If you've ever worked with any workflow engine, doing it with agents is largely the same. If you were writing some automation that used a browser, how would you handle recovery for any given step of your workflow? It depends on what you're doing, the specifics of the web app you're interfacing with, etc.
Contrived example, but let's say you're sending an email. Load the page, click the button, enter text in the various fields, etc. Since there's no side effect of consequence until you hit send, you could just make sure your failures clean up drafts, and replaying the whole thing is safe.
I've admittedly done very little browser automation like this though, mainly I've just called APIs, created and uploaded files, done db operations, normal dev stuff.
I'd expect RPA platforms to be way ahead on agent automation like this, I haven't kept with any of them though. If they aren't, real missed opportunity for them.
So rather than saying “open chrome, go to this page, click next page 5 times”, it would be something like “chrome is running; url is X; url is X/page/1; url is X/page/2” etc. Ansible, basically.
Other than that, you could just replay bash tool calls. That’s full of holes, though. Anything that relies on “date” will return different stuff, and if you try checkpointing the system time then TLS breaks due to timestamp differences.
If you wanted to go absolutely wild, some hypervisors can checkpoint the memory of a running VM and revert back to a prior version memory and all. I can’t imagine a way to make money off that (you’d be writing gigs of data per checkpoint), but I suppose it’s technically possible.
Woah, that big of a difference when it comes to token counting?
https://openrouter.ai/blog/insights/opus-47-tokenizer-analys...
I don't want to quibble on definitions and semantics here but by this definition, there wouldn't be a single harness out there that I can see except those that include paid cloud storage like Dots.
Or my laptop crashes, ugh.
Yes, if I could ssh into a random server it'd be fine. But I can't.
I do use a git as pi memory with commit history and a worklog. Restarting pi just continues from the worklog + last commit/uncommitted changes.
Durable ... is different but I definitely will try it out.
``` const SyncToPostgres = defineTask<{ entryId: string }, { phase: "send" }, void>({ kind: "app.sync-postgres", version: 1, initial: () => ({ phase: "send" }), phases: { send: async (task, runtime, context) => { const entry = await runtime.read(/* the entry */); await postgres.upsert("messages", { id: task.input.entryId, ...entry }); // idempotent by id await runtime.commit(() => ({ status: "terminal", outcome: { status: "completed" } }), context); }, }, });
await root.commit(async (tx) => {
const id = await tx.entry(AssistantEntry, answer);
await tx.createTask(SyncToPostgres, { entryId: id }, { ownership: { kind: "conversation" },
background: true });
}, context);
```The commit on the root conversation picks out the last agent answer id from the transcript, and durably schedules a task that then syncs it to postgres. inside the task, you fetch the answer by id and send it over to postgres indempotently.
What's missing here is sugar, basically a hook that runs inside each commit so the outbox write is atomic with the state change, with ordered delivery, and possibly a durable change feed with cursors.
Thanks for the input!
also, I see most of the durability promise comes from persisting JSON documents locally and minimizing the amount of context/data kept in-memory, even during SQLite mode. while this makes sense, my own experiments with a process that relied on a JSONL-based event store have led me to prefer keeping things in-memory to avoid all the friction with I/O.. am I crazy for preferring just a straight .db file being persisted?
If you add e.g. bash as a forced default tool, then someone can't come up with an extension called "sandboxed-bash", which internally runs the sandboxing logic and then delegates back to the bash tool.
By baking in your specific personal use cases you have made your software tool useless to the vast majority of people on the planet. Some of those people might decide to go ahead and use your software anyway and then run into massive headaches along the way and pretend those headaches aren't real, but that doesn't change the fact that the software design is incredibly poorly though out.
A coding agent doesn't necessarily need to write files. A review agent can just read the code, maybe it doesn't even read files on disk, maybe it just looks at a code diff on github and then posts a line by line comment. It does not need bash or node or whatever default tool you think is cute. It needs the tools I give to it and if it uses only the tools I give it, then I don't have to babysit it. If you let it run bash or node just to be cute, I have to babysit your agent harness. Is that so hard to understand?
If I need 100 different agent types, and they all have bash or node and there is a risk of them using bash or node when I only want it to use exactly the tools I want it to use, then why the hell would I choose your software? I wouldn't. I don't want to use it. It is completely illogical. Some people want to run agents as if they are microservices. Yes, that's me. I don't want to babysit every single microservice. You guys want to build the ultimate agent monolith and then call it minimal.
I got burned so I'm going to write my own harness anyway. Have a nice day.
Sounds like a bug under a false assumption. Just having an ID cannot alone guarantee exactly once semantics AFAIK.
The harness is built on top of the Pi SDK. I initially used Codex, but Pi seems more hackable, and I like that it’s vendor-agnostic by default.
Running it on Kubernetes works, but dealing with the JSONL session files and making sure sessions survive pod interruptions adds some complexity. I’m using DBOS for that right now, which works well, although it still feels like overkill.
This came at just the right time. I’m looking forward to removing the pieces I no longer need and simplifying the architecture. Thanks Pi team!
When I am using Pi to write extensions for Pi, I feel better running Pi wrapped in a separate os-level sandbox. I guess Pi could do it all, but I am content with how it is.
These should be decoupled.
Maybe I need nono in one context and smolvm in another or both.
I would not want to trust the harness to self policy.
However, for ease of use, it is nice for harnesses to by default run with sane and safe sandboxing setup. Then give the option to disable them.
Isn't it better if the tool is sandbox agnostic and you as the developer / integrator choose what's best for your use case? There are several levels of sandbxing, with many degrees of "freedom", so it would be really hard/confusing/overly-complex to build something ootb that suits everyone, no?
Within a session, you can give each conversation its own sandbox, based on your application's needs and policies.