Remote Code Execution as a Service
earthly.dev
earthly.dev
We built a service that executes arbitrary user-submitted code. An RCE service. It's the thing you're not supposed to build, but we had to do it.
Running arbitrary code means containers weren't a good fit ( container breakouts happen), so we are spinning up and down ec2 instances. This means we have actual infrastructure as code (i.e. not just piles of terraform but go code running in a service that spins up and down VMs based on API calls).
The service spins up and down EC2 instances based on user requests and executes user-submitted build scripts inside them.
It's not the standard web service we were used to building, so we thought we'd write it up and share it with anyone interested.
One cool thing we learned was how quickly you can Hibernate and wake up x86 EC2 instances. That ended up being a game-changer for us.
Corey and Brandon did the building, I'm mainly just the person who wrote things down, but hopefully, people find this interesting.
The earthly backend runs on a modified buildkit so it is running the arbitrary code in a container, but it's also in its own VM. This was simpler then firecracker to get started but turned out to have pretty good performance and alright cost once we started suspending things.
It's possible to mitigate/reduce them for sure, with appropriate hardening, but the Linux kernel is still quite a big attack surface.
And you can consider using gVisor to minimize container breakouts to a great extent.
gVisor was considered but so far it looks like the next iteration with be using firecracker vms. Our backend is buildkit and it can't run in gvisor containers without some work.
How do you run firecracker?
So it wasn't so much disqualified as decided it wouldn't be the v1 solution. We wanted to get something out and get people on it and get feedback. So weren't afraid to spend some more computer dollars to do so.
Could you talk more about it? Are you keeping a cache of hibernated EC2 instances and re-launching them per request? What sort of relaunch latency profile do you see as function of instance memory size?
That instance is just sitting around waiting for gRPC requests that tell it to run another build. If it's idle for 30 minutes, it hibernates and then if another call comes back in a gRPC proxy wakes it back up.
I don't know if the wake up time increases per the size of the cache in memory, I can check with Brandon but its much faster starting up an instance cold, mainly because buildkit is designed for throughput and not a quick startup.
There are more details in the blog.
if a build 1 happens to install a specific libc, do you un-install that libc before running build 2?
if you just say that stuff is the responsibility of the user, okay, but then the artifacts produced by this system aren't deterministic, which seems like a problem?
Buildkit runs the builds in runC, so basically containers are used to keep things deterministic, but the buildkit backend isn't shared, each is in own EC2 instance.
Converting Docker images to run in VMs instead of containers.
Source: ran a container based RCE service that ran millions of arbitrary workloads per month. We had sophisticated network and system anomaly detection, high priced pentesters etc and never had a breakout.
Would "never detected a breakout" be better wording? :)
I assume GP wrote that in order to say that they have a high confidence that they never actually had a breakout.
You are technically correct. But your logic applies to everything. Is the isolation provided by VMs good enough? Is airgapping enough to prevent breakout?
There are many things that factor in when you decide what's reasonable. Some are first principle arguments (containers use the same kernel as the host, the kernel has a large surface area, ...). Others are statistical arguments: there have been past breakouts with this stack, it's thus reasonable to expect more in the future, ...
We found that privileged is a pretty big hammer and thought we needed it too but we found ways to give us the functionality we needed without all the extra stuff we didn't need the privileged brings in.
if user code can modify the VM, and VMs are stateful, then how do you ensure hermetic builds?
Buildkit runs the builds in runC, so basically containers are used to keep things deterministic, but the buildkit backend isn't shared, each is in own EC2 instance.
So if you could break out or access cache entries that weren't correct you would have found a way to break reproducibility, but not access anything you shouldn't.
> but the buildkit backend isn't shared, each is in own EC2 instance
it isn't shared between different customers, but it is shared between different builds for the same customer, right?
I myself did try to run buildkit in a Lambda as I think that would be low cost option. But I found it you couldn't make gRPC calls against a lambda and that is a hard requirement for us.
Still, if one were to attempt this exercise -- one path might be via extremely small emulated instruction set in an extremely tiny and open source (much easier to audit by end users) virtual machine where the emulated instruction set is NOT the same as the one on the host machine (i.e., RISC-V emulated instructions on x86 or ARM host...).
Would such an emulator be slower than native host execution?
Probably, possibly many times slower...
But, what might be lost in speed -- might be made up for in the ability to audit (at least better) for security concerns...
So in summation, is it a good idea?
Possibly -- but it needs that extra rigorous security guarantee which is very difficult to get exactly right in this day and age...
Also -- supposing that this business model was completely viable from a security perspective -- the next step to building such a system might be to build an auction site for CPU cycles (or compute ability/compute time) -- sort of like an "EBay for highly distributed but sandboxed code execution" (clients bid on renting compute resources, end users bid on selling theirs) which would be an interesting business model, IMHO...
Related: SETI@home: https://en.wikipedia.org/wiki/SETI@home
In Commonwealth saga from Peter F. Hamilton, people have access to nearby computing nodes like we have currently to free wifi, for accelerating any tasks they would need (like personal assistants, called e-butlers). That's not 100% secure and is sometimes used for nefarious purposes (and plot devices), but it's just accepted as a nuisance like other things that can be weaponised in our own reality (like cars).
Note: All the interesting stuff is in /opt/appfs
Spawned a program on every single node that wrote "yes" into the file /etc/am_i_compromised..