Workerd: Open-source Cloudflare workers runtime
blog.cloudflare.com
blog.cloudflare.com
I've dug a bit on the OSS repo [2], and here are some of the findings I got:
1. They are using a forked version of V8 with very small/clean set of changes (the diff with the changes is really minimal. Not sure if this makes sense, but it might be interesting to see if changes could be added upstream)
2. It's coded mainly in C++, with Bazel for the builds and a mix between CapNProto and Protocol Buffers for schema
3. They worked deeply the Developer Experience, so as developer you can start using `workerd` with pre-build binaries via NPM/NPX (`npx workerd`). Still a few rough edges around but things are looking promising [3].
4. I couldn't find the implementation anywhere of the KV store. Is it possible to use it? How it will work?
Open sourcing the engine, mixed with their recent announcement on a $1.25B investment fund for startups built on workers [4] I think will propel a lot the Wasm ecosystem... which I'm quite excited about![1] https://twitter.com/KentonVarda/status/1523666343412654081
[2] https://github.com/cloudflare/workerd
1. We float a couple patches on V8 but we keep it very up-to-date. Generally by the time a version of V8 hits Chrome Stable, we're already using it. I would indeed like to upstream the patches, which should be easier now that everything is open.
2. FWIW the one protobuf file in the repo is from Jaeger, a third-party tracing framework. We actually aim to factor that out into something more generalized, which may eliminate the protobuf dependency. Workerd is very much based on Cap'n Proto for most things (and the Workers team does most of the maintenance of Cap'n Proto itself these days).
3. Yep, one of our main goals was to bake this into our `wrangler` CLI tooling to make local testing easier. I'd also like to ship distro packages once things are a bit more stable.
4. The code includes the ability to configure KV bindings, so you can run code based on Workers KV. But, the binding is a thin API wrapper over an HTTP GET/PUT protocol. You can in principle put whatever you want behind that, and we plan to provide an implementation that stores to disk soon. The actual Workers KV implementation wouldn't really be that interesting to release as it mostly just pushes data over to some commodity storage back-ends in a way that's specific to our network.
A team on workerd? I'm sure there is a lot of work involved in running the massive Workers platform, but curiously enough, the workerd git repo itself has relatively few commits from other engs... just where is the team? ;)
https://github.com/cloudflare/workerd/commit/a265d5c22bf2a6b...
> Cloudflare is not providing any funding or making any funding decisions, and there is no guarantee that any particular company will receive funding through the program. All funding decisions will be made by the venture capital firms that participate in the program. Cloudflare is not a registered broker-dealer, investment adviser, or other similar intermediary.
So it’s not a fund. Cloudflare just has an agreement that a set of venture funds will set aside capital for companies that leverage workers and in return Cloudflare will provide deals / resources to said startups.
The reason I bring this up is because this is significantly less helpful to founders building on their platform than simply having a strategic fund of their own. The tl;dr on why is because it’s an additional hoop to jump through to get to an investment decision.
As a founder that has raised a significant amount of strategic capital and operates in the same space as CF Workers, if you want companies built on your platform funded, optimize for fast process. I worry the dealflow will actually work in the wrong direction here, with these firms simply telling startups, “we’ll give you some CF credits if you build stuff using workers — helps us hit our promised allocation with Cloudflare.” Which might be the intent, but running a strategic fund is probably more founder-friendly and may create better outcomes.
Our small tech shop has been using Workers since Jan 2020. Leaving behind the monstrous and intimidating world of AWS, I simply couldn't contain my excitement, even back then. In the words of Eminem, Ease over these beats and be so breezy; Jesus, how can shit be so easy?!
Today, the entire platform around Workers is nothing short of astonishing. With Cloudflare's clout, Kenton and his team have utterly transformed serverless computing, staying true to the lessons-learnt from running sandstorm [0] way back when [1]!
It is almost comical (or diabolical, depending) that none of the Big3 have bothered competing with them. It is only the upstarts like Bun, Deno, SecondState, Kalix (Lightbend) and others that are keeping Cloudflare honest.
I am originally from the PHP/Laravel world, recently had to quickly develop a very simple (5 endpoints) REST API and decided to give CF workers & KV storage a try. In an hour the API was running in production…
CF is really, really good at many things.
AWS uses/maintains Firecracker which meets criteria to be considered secure but has a 300ms start. I've been confused how Cloudflare can sandbox things safely without the same cost
so the answer is, a lot of things: * v8 isolates
* private processes for debugging/inspection
* "cordoned" V8 runtimes organized by level of trust (guards against V8 zero-days for one)
* linux namespacing as an outer sandbox layer, blocking fs and network access
* explicitly allowed capability-based security ...etc.
Amazed by what OP and the Workers team have done over the years. Took a while for us to get used to the Workers paradigm. But once we did, feature velocity has been great.
Last wish list item: a Postgres service on Workers (D2?) becoming available in the not too distant future
Not sure where Cloudflare's with relational database connectors [0], but the next best thing is MySQL via Workers: https://news.ycombinator.com/item?id=32511577
[1] https://blog.cloudflare.com/relational-database-connectors/ [2] https://blog.cloudflare.com/whats-new-with-d1/
We've mostly been using the GA plugin for use with beta users. Had completely forgotten about the need to verify the app with Google. Will do that this week.
https://blog.cloudflare.com/introducing-workers-durable-obje...
Durable Objects aren't exactly the same as Erlang actors (which is why we don't usually call them "actors" -- we found it led to some confusion), but at some level it's a similar concept.
You can easily have millions of Durable Objects in a Workers-based application.
Currently workerd has partial support for Durable Objects -- they work but data is stored in-memory only (so... not actually "durable"). The storage we use in Cloudflare Workers couldn't easily be decoupled from our network, but we'll be adding an alternative storage implementation to workerd soon.
They remind me of this article: http://ithare.com/scaling-stateful-objects/
One critical thing that I couldn't find info about is what reliability to expect from the persistent storage for a Durable Object. Is it more or less a write to a single disk without redundancy? If there's redundancy, to what degree? Essentially, how much would you need to build on your own in terms of replication for a production scenario if using Durable Objects as primary storage?
I get that they can use other, separate storage if necessary, but either way it seems like an important consideration when designing a system on top of them.
The fact that there is one unique instance of the object live at a time (and therefore only one client for the distributed storage) lets us do a lot of tricks to hide the storage latency.
https://blog.cloudflare.com/durable-objects-easy-fast-correc...
However, I don't entirely get it. If every machine runs every service with library-like overhead for calls across them, how is that different from a monolith that links in libraries? It just seems more complex for no benefit.
I guess it's not a problem I've encountered so far. Perhaps I just haven't worked in a company big enough? Only on the scale of hundreds so far.
Branching, agreeing on interfaces and tests have been sufficient to allow linking, in my experience.
For example, a client wants to write a workflow but has a custom step that integrates with their system only. They can write the code and run as part of our workflow without us worrying about them writing something nefarious and breaking out of the sandbox.
That said, the workerd github says that this can't (yet?) be relied on to be a security sandbox.
workerd ensures that a Worker cannot access protected resources by accident or negligent programming. What it can't protect against (on its own) is intentional exploitation of security bugs.
If your clients are people whom you've spoken to face-to-face and signed contracts with, then you can probably trust that they won't exploit a V8 zero-day to break the workerd sandbox, because if they did, it would clearly be a crime and you could prosecute them.
If your clients are random people who can sign up on the internet without talking to anybody, then they may be able to evade such identification and therefore prosecution. In that case you need a stronger sandbox to contain them.
But still isn’t the Wild West.
It would likely be more of a “this person wrote crappy code that’s bringing down everyone else’s worker” or “this client wants to know if we have proper guards on their data”
A regular javascript module can be deployed independently. And have its own version of dependencies. I'm just not seeing what this really adds over existing widely available, well understood options.
I guess protection from other teams uncontrolled monkey-patching is something (I don't think there's probably any legitimate reason to monkey patch something yourself in normal production -- why wouldn't you patch or replace the problematic thing?), but there might be other ways to deal with that short of a new runtime environment.
I'm a little worried truly independent updating... imagine starting a request with one version of the auth module but finishing it with another. Not sure people really want to have to account for that. (Maybe that's accounted for, so maybe an unwarranted worry.)
I understand why I'd want to do edge compute: it lets me shift work (in some cases at least) from my backend to the CDN, so my backend gets less traffic and users (possibly) get responses faster since the workers run on CDN edge nodes.
But obviously if I'm running workerd myself, I'm not doing edge compute (unless I'm also operating my own CDN). And if I'm not doing edge compute, just deploying code or a container to a server seems easier. Is it for the case where you want to do serverless stuff in a runtime you host yourself?
But the other reason you might develop on workerd is because you want the option to deploy to Cloudflare later. The advantage of Cloudflare isn't just that it's "edge" and therefore closer to clients -- it's also just a lot easier to deploy code on Workers a lot of the time than it is to manage servers on traditional cloud providers.
And cloudflare making it easier to migrate away, also makes it easier to migrate to, since you aren't quite so stuck.
Could be some things aren't allowed to run in the cloud for gdpr, privacy reasons or whatever
I fear that, every worker could end in spaghetti Code and nobody remembers why they used this or that and where it was saved.
You could follow the path of the request, but I see some delevopers going crazy with workers.
I use workers too, I love them.
We support typescript (deno), go and python. Short of those languages becoming 100% transpilable to wasm, we will have to fallback to use firecracker or keep our current system implemented from scratch using nsjail for sandboxing and a fork for each process.
But as far as JavaScript/Wasm, that's a pretty inherent design feature to the Workers architecture. Cloudflare will probably provide containers eventually but that'll likely present itself as more of an add-on than a core part of the platform.
The problem is, containers are just too heavy to support the nanoservices model as described in the blog post. JavaScript and Wasm are really the only viable options we have for something lighter-weight that meets our isolation needs. Other language VMs are generally not designed nor battle-tested for this sort of fine-grained isolation use case.
Does this mean that TCP workers would be exposed to network level attacks or would it use transit/spectrum? If it turns out to be protected; I'd say there would be little to no reason to use Spectrum unless the pricing turns out to be atrocious for long lived connections (which is kind of the point of having TCP workers in the first place).
I hope I did not come out as rude; I'm genuinely curious about what's the plan behind all of this.
Edit: I pointed out there would be no use for spectrum since one could "easily" build a reverse proxy with a tcp worker.
https://github.com/cloudflare/workerd/commit/a265d5c22bf2a6b...
We do have some plans around making WebSockets more efficient by allowing the application to shut down while keeping idle WebSockets open, and start back up on-demand when a message arrives.
FWIW contrary to the note on that page, WebSocket pricing is documented along with Durable Objects here:
https://developers.cloudflare.com/workers/platform/pricing/#...
> If 100 Durable Objects each had 100 WebSocket connections established to each of them which sent approximately one message a minute for a month, the estimated cost in a month would be, if the messages overlapped so that the Objects were actually active for half the month:
I remember looking at that page before and finding these two lines confusing. Based on the first line, it seems that the objects would be considered active for the entire month, whereas the second one raises the possibility that a websocket can be connected to a durable object without it being considered active.
> A WebSocket being connected to the Durable Object counts as the Object being active.
Please make noise when that rolls out! I was literally wishing for that yesterday.
I want to use Cloudflare Workers to quickly connect p2p WebRTC clients, but when a new peer joins there needs to be open WebSockets to send connection-setup messages to each peer in the room. But with how that works / is billed today it really doesn’t make sense to have a bunch of idling WebSockets.
An alternative is having each client poll the worker continuously to check for new messages, but that introduces a bunch of latency during the WebRTC connection negotiation process.
In order to support efficient WebSocket idling, we probably need to have some ability to respond to application-level pings without actually executing JavaScript. TBH I'm not exactly sure how this will work yet.
That would mean maintaining "chatroom" state in memory wouldn't be enough, the DO worker would have to always write it to the durable store, to be safe to unload.
You'd have some sort of a "WebSocket connection ID" that could be used as key for storing state, and when the next message comes in, the worker would use that key to fetch state.
EDIT: And the ability to "rejuvenate" other WebSocket connections from their IDs, so you can e.g. resend a chat room message to all subscribers.
Using Cap'n Proto instead of some bloat <3
Abstractly speaking, I expect we need to build some workflow tooling around the nanoservices concept, since somehow all your teams' nanoservices eventually need to land on one machine where one workerd instance loads them all. But should we do that by updating whole container images, or distributing nanoservice updates to existing containers, or...?
This would be a great discussion topic for the "discussions" tab in GitHub. :)
Can anyone give some examples of this? I’m not sure I’m following exactly here. Is this referring to isomorphic (JS) API’s?
[edit:typo]
How many non-javascript folks use this? What's better than using fly.io?
It's possible that Cloudflare and Fastly users are different enough that different approaches make sense.
We are trying to figure out how to move more of the API parts of runtime to something not c++ without significantly negatively impacting memory / CPU (safety, correctness, speed of development, etc). It’s an interesting problem to try to do at Cloudflare scale.
Currently there is no way to test for cloudflare-esque failures. For instance, what does my worker do when the kv lookup fails or another cloudflare infrastructure related promise isn't fulfilled? I think actually running the true worker runtime in a testing environment is key.
(In my opinion objectively the worst possible choice because + requires URL encoding ;))
I periodically get crap for it but so far it hasn't really broken anything.
The biggest annoyance is exactly what you point out -- the GitHub URLs look ugly because they %-encode the plus signs. Though it turns out you can manually un-encode them and it works fine.
Do you have a reference to the advice to disable default rules? It sounds counterproductive to me.
Jokes aside though, the only major complaint people have is about .RECIPEPREFIX, which I also don't follow. I didn't read too many comments, but the top voted comments do not appear to take issue with the advice to disable built-in rules.
I'd be surprised, since it is far from the first time I've heard that advice, and it never seemed very controversial.
So I used the actual name of the language, and it worked fine. :)
this basically becomes the first large scale deployment of wasm in the world right ?
can someone talk about the engineering choices here. why this ? why not rust/java/c++...any of that ? a people perspective.
I'm not sure what to explain here? It's a runtime based on V8 to run JS and WASM code. Used by Cloudflare to run on edge. Why not just rust/java/c++? Well, good luck running that in multi-tenant environment.
First, all 3 allow far too much: uncontrolled file system access, uncontrolled network access.
Second, customer's code cannot be trusted. Running unknown code on your server leaves with huge attack surface.
Third, isolation to deal with first two is expensive: resources and cold-start.
Third+, the point of running code on edge is for it to be quick. No one will use it if your cold-start time is higher network round-trip, and you're not going to get paid a lot if you have to run customer code even if not in use.
Fourth, running WASM allows unified runtime without any dependencies:
- Java, which JVM do you target? What language version do you target? How often you upgrade? Do you run one JVM per tenant? How do you deal with cold-start times?
- Rust/C++ which architecture your servers use? What if you want to switch to ARM? What if it's mixed? Which version of ARM it is? Can I use AVX-512?
- WASM: WASM is WASM. You don't care if customer compiled rust or C++ or Go to get that wasm.
Just wondering if it makes sense to make that investment. Some of these problems u will encounter when WASM gains a GC anyway. This is one of the big reasons why Graal doesnt target wasm backend. So this - one runtime/gc vs multiple is something that is a problem ull have to solve anyways.
The original JVM was intended to provide secure sandboxing for "applets" but that use case failed in the market, and instead JVM became focused on use cases where isolation isn't important. Years later as security research became stronger, all sorts of holes were found in the JVM's sandbox. Presumably those holes were fixed, but if we take JVM and use it in a security-critical environment again, probably a bunch of new holes get found pretty quick?
(I don't know anything about Graal specifically, I guess it's an alternative JVM? But how much security research has it had?)
Whereas V8 has had constant security research and real attacks over the course of about 15 years. It's certainly not bullet-proof either but we have a pretty good idea of just how much of a risk it is.
Again, catch 22, there's seemingly no path forward for these VMs to prove themselves. I don't know what to do about that, but that's where we find ourselves.
GraalVM is interesting, and it has isolates and from my memory it has plenty of sandboxing features like limiting network and file system access. However, it has some big downsides: it's Oracle, people still associate it with your regular JVM and all its downsides. Yeah, it can match v8 in performance if you let it warm up (it takes longer for it to warm up compared to V8.) and it's even further from what CF does in terms of compatibility with existing ecosystem.
CloudFlare workers have single digit cold-start time[2]. GraalVM is...200ms and that is considered blazing fast by JVM standards. Takes longer and slower at runtime and probably more memory consumed.
Remember, we're not talking about long-running services, we're talking about short-lived request handlers in JavaScript that run on edge. Pricing for workers is per request. It's in CF interests to serve customers requests as fast as possible. Also remember that customers want to write JavaScript.
Why is Oracle an issue? GraalVM is under GPLv2 which does not cover patents. Some performance and security features locked in Enterprise version. In addition, first production-ready version was in 2019, CF already launched workers by then.
[1]: Last commit 2018. I doubt this is a feature complete software...
[2]: https://blog.cloudflare.com/eliminating-cold-starts-with-clo...
We've always designed Workers around open standard APIs, so even without this, Workers-based code should not be too hard to port to competing runtimes. But, this makes it much easier to move your code anywhere, with bug-for-bug compatibility.
You can't build a successful developer ecosystem that has only one provider. We know that. We'd rather people choose Cloudflare for our ability to instantly and effortlessly deploy your application to hundreds of points of presence around the world, at lower cost than anyone else. :)
I'm curious on the use cases for self hosting? Ie i have a heavy interest in self hosting. I also find workerd very interesting. Though i'm struggling to think of reasons why i'd build my apps around workerd instead of just a process.
If i was dedicated to WASM, Perhaps workerd might be a better solution than wasmer, unsure.
Reasons you'd write first party code targeting workerd would be a different question.
---
Cloudflare... uses... workerd, but adds many additional layers of security on top to harden against such bugs. However, these measures are closely tied to our particular environment... and cannot be packaged up in a reusable way.
Durable Objects are currently supported only in a mode that uses in-memory storage -- i.e., not actually "durable"... Cloudflare's internal implementation of this is heavily tied to the specifics of Cloudflare's network, so a new implementation needs to be developed for public consumption.
...the runtime parts I was most curious about couldn't be open-sourced :(
https://github.com/cloudflare/workerd/blob/main/src/workerd/...
We definitely need to provide more higher-level documentation around this, which we're working on, but if you're patient enough to read the schemas then everything is there. :)
To act as a proxy, you would define an inbound socket that points to a service of type `ExternalServer`. There are config features that let you specify that you want to use proxy protocol (where the HTTP request line is a full URL) vs. host protocol (where it's just a path, and there's a separate Host header), corresponding to forward and reverse proxies.
Next you'll probably want to add some Worker logic between the inbound and outbound. For this you'd define a second service, of type `Worker`, which defines a binding pointing to the `ExternalServer` service. The Worker can thus send requests to the ExternalServer. Then you have your incoming socket deliver requests to the Worker. Voila, Workers-based programmable proxy.
Again, I know we definitely need to write up a lot more high-level documentation and tutorials on all this. That will come soon...
Right now, you can't use most NPM modules on CF Workers since this is a custom runtime which is also pretty barebones (probably by design).
Deno in comparison offers a rich standard library which solves plenty, but they are also working on a compat layer for Node modules. And the Deno ecosystem is also growing, probably because you can run it anywhere.
Anyone else thinks the name is kinda weird? After reading the article it makes sense that it's "worker d" but I first read it as "work erd" or a single word like "worked" but with a rogue R in the middle.
Honestly, it's a boring name. We had a big debate over naming with a lot of interesting ideas, but ultimately every idea was hated by someone. No one loves workerd, but no one hates it, so here we are! Oh well.
Obviously I'm not a Unix guy and looking at the downvotes on my previous comment, people seem bothered by that :)