Celld: Self-hosted, distributed Durable Objects
github.com
github.com
> Pull requests are disabled. Coding agents make it too easy to send a large,
> low-context change that costs maintainers more time than it saves.
> Thoughtful contributions are welcome; please understand the code,
> keep the patch focused, and respect the review time you are asking for.
>
> Send a git format-patch attachment to ...> At this time, we are not seeking outside contribution.
> AI has made writing code easy. The hard part, today, is not writing the code, but reviewing it, making sure quality stays high, and keeping the product coherent. In that light, unfortunately, external code contributions are "donating" the easy part of the job, while creating more of the hard work.
Feels very weird, but is logically sound: owners know exactly what they want and so they can work with Claude et al to iterate on features faster than with most drive-by contributors.
It reads like an "end of an era" but I imagine the steady state will be somewhere in the middle: high trust, high context contributors will still be able to contribute meaningful work.
CloudFlare OS: https://blog.cloudflare.com/cloudflare-os/ https://news.ycombinator.com/item?id=49182996
And yes, celld is isolates! The very lightweight within v8 isolation boundary. Hence the "very low idle cost". The deno team rolled their own new runtime! This one without deno_core! Shout out to this excellent 2022 post: https://deno.com/blog/roll-your-own-javascript-runtime https://news.ycombinator.com/item?id=35819990 https://news.ycombinator.com/item?id=35819990
It's super super exciting having an age where many pieces of software can all run with low profiles. Very timely, again, for the unbelievably order of magnitude (and thats 10 to the not 2 to the, for you fellow computer nerds) ish ram crunch we are in. Next up, at some point, ideally we get some v8-like runtimes where we can share libraries across multiple isolates! Separating the shared code from the shared data, so we have multiple instances of the library, seems much harder. (It feels like wasm has some/much promise here and I eagerly await clearer signals that code-sharing while sandboxing is indeed possible)
I'm probably missing your point, but aren't wasm memories achieving exactly this? Just an example from wasmtime: https://docs.wasmtime.dev/examples-multimemory.html
So happy to see support for running durable objects outside of one provider. Upvoted.
The "durable object" concept has been repeatably demonstrated to be a valuable abstraction.
"Each object is its own SQLite database, addressed by name and replicated to an S3-compatible bucket you own" -- this concept can take you a long way, both in its power and simplicity.
There was also Jamsocket and Plane, though it looks like they've shut down after acquisition: https://github.com/jamsocket/plane
This seems to be the first one which is providing drop-in compatibility for Cloudflare's JS-side APIs, including JSRPC etc.
https://www.microsoft.com/en-us/research/wp-content/uploads/...
celld is the full distributed system (albeit single tenant). It distributes DOs (cells) across any number of VMs. The databases for each DO are in object storage with RPO=0 guarantee.
https://github.com/cloudflare/workerd/pull/6780
But honestly I love that there are multiple implementations now.
ah well, we have celld!
The goal of this new design is to scale to a cluster while being operationally very easy to set up. Ideal for self-hosting.
NFSv4 is an easy first step, convenient because it's broadly understood, has many implementations, and requires no client libraries. I could imagine a follow-up to support LiteFS instead of NFS could make a lot of sense, though it'll get more complicated.
But yes, celld is definitely ahead of us here. No doubt about that.
I've always wondered what it is Cloudflare does to make DO's work but this specific thing doesn't seem explained anywhere
Dual-licensed under the GNU AGPL v3 (fully featured, for open source use)
and a commercial license.
but the commercial license link https://www.zerofs.net/licensing is a 404 pageMore practical example (serverless WebSockets) https://youtu.be/FgWVoryZ8PU
It worked really well! Excited to see more options outside of Cloudflare.
celld's README states "Object-storage compare-and-swap ensures that exactly one node owns a cell at a time," but I'm skeptical this actually holds at the point where data is written to storage - I had an AI skim through the code with me, and the actual segment writes looked like plain, unconditional PUTs with no epoch check.
In many cases, I think using Cloudflare's Durable Objects is probably the right call instead.
i would prefer a thing that was more self-contained, not dependent on a black box service layer underneath.
much apologies if i just have a poor understanding.
It's like fixing all the problems with democracy by putting everyone in charge of their own 1-person election.
Short writeup: https://crabmusket.net/2024/durable-execution-versus-session...
Summary:
I think the way to decide is, do you want to program with "objects" or "processes"?
I'd use Durable Objects (or what I called Session Backends in my article, after Jamsocket) if I wanted to model an "entity" that lasts indefinitely (e.g. a Figma doc, a user, a concert/event). I'd use a temporal/restate function for something that has a linear sequence of events and eventually comes to an end.
This isn't a hard rule. You can get the same outcome out of both technologies. (For example, Cloudflare built Workflows on top of Durable Objects; Rivet did the reverse.) You can use either one to implement the other, but you're going to be going out of your way based on what APIs are provided.
It's just making it way easier to spam maintainers with code.
Good contributors are probably worth even more now.
cloudflare gives you that instant geo-sync across the world that is hard to beat
with Celld do i need to buy bare metal in major continents
thats hard to beat
on the other hand, deno land folks stole from (used open source from) the best. the coordination layer/control-plane is all S3 CAS of dumb json files, which is a fantastic common-mode infra requirement for most orgs anyways, & perhaps a durable control-plane substrate you'd feel comfortable having someone else run (such as aws or others), while you run the data-plane (workers) yourself. and then for the durable objects themselves, they used litestream, which is a pretty top pick, excellent way to get radical distribution (but actual topology not included, some assembly required)! https://hn.algolia.com/?q=litestream
there's no reason this wouldn't run fine on most hosting, you definitely don't need bare metal. the virtues of v8 sandboxing / isolates! no need for vm's at all, no nested vm difficulties if you are trying to host on a shared host! but if you're asking questions like this, i want to again point you back to my top point.
So where it's overkill, this is where Celld is a good fit. It's a 100% open play, despite mentioning three hosted APIs: S3, Cloudflare Workers, and Durable Objects. The thing is, that Deno tried making similar ones. Deno Deploy is analagous to S3, and Deno KV is analagous to Durable Objects. However, with Celld the only dependencies are being able to run a binary and some S3-compatible storage. Cloudflare provides S3 compatible storage, but this README doesn't mention that, except in an example. Perhaps Ryan is leaving the door open for Deno to be acquired by Cloudflare.
FWIW here's an old Deno blog post comparing Durable Objects to other stuff: https://denoland.medium.com/deno-kv-vs-cloudflare-workers-kv...
the stuff is great when it works but Durable Objects can be quite expensive. whenever i get too excited about em all it takes is a little time trying to price it out to calm me down.
> do i need to buy bare metal
my first idea would be to run celld on AWS Kubernetes deployed to local zones https://docs.aws.amazon.com/eks/latest/userguide/auto-local-...
it’s not “region: earth” like cloudflare but perhaps worth the trade off.
... but that S3 _is_ the control plane and consensus layer, no? You're just pushing this down the stack to whoever runs that S3 clone.