Using Firecracker and Go to run short, untrusted code execution jobs (2021)
stanislas.blog
stanislas.blog
The scripts would be long-running, but don't require much computational power. Just a few arithmetic operations on an array every second. The array is being fed in via WebSocket.
Thank you for linking gVisor, went through the docs real quick and it looks very promising for my use case.
Using Firecracker directly is pretty straightforward: https://jvns.ca/blog/2021/01/23/firecracker--start-a-vm-in-l...
gVisor gives you all the container tooling, which may or may not be useful.
And, because I'm a shill, we actually shipped an API specifically for this kind of use case. So if you'd rather not build it all yourself, we can help: https://fly.io/blog/fly-machines/
https://news.ycombinator.com/item?id=32289979 (70 points, 16 comments)
and
https://news.ycombinator.com/item?id=32287798
Both links include hot takes from a tech lead for CF Workers (@kentonv).
Happy to expand more of my experience of making this work at scale.
Docker + heavily restricted user + firewalls.. seems to get you much of the way there. I am aware that some work was done back in the pre-Docker day with Ruby's online sandbox to neuter Ruby's ability to make certain syscalls, but I imagine Docker, eBPF, or even using WebAssembly makes it a lot easier now.
gVisor is kind of Linux running in Linux. More insolation than containers; less overhead than VMs (but less isolation, of course).
> less overhead than VMs
Make sure to benchmark your workload first -- gVisor's I/O subsystem is a lot slower than the Linux kernel's, so a VM can be materially faster if you're doing a lot of filesystem operations or file I/O.One of the systems I built at a former employer supported both gVisor and Firecracker for isolation, and the gVisor version was 10-50x slower for a specific class of workload that did ~millions of stat() calls at startup.
One thing to bear in mind is that these sites use super-paranoid security because it has been proved time and time again that it is necessary. I wouldn't look at any particular solution for running arbitrary code from a user and assume that it's actually 100% correct. I think this can help remove some of the mystery of how they do it, which is that there is very likely some way in which they actually aren't doing it. Once you remove that idea from the possibility space, the ways it is done start making much more sense. (And the idea becomes much more scary.)
Does anyone know about this?
Theoretically, it should be possible to do it all in the browser, just a lot more work given the state of things now.
* https://github.com/kripken/llvm-js
* https://github.com/tbfleming/cib
* https://github.com/jprendes/emception
LLVM just does pure computation, really, so it's not hard to port to wasm - much simpler than say Python (which has also been ported several times). The only challenges with LLVM are the build system (which has self-execution), working around some issues like clang wanting to open a subprocess, and adding some ifdefs.
The history here has been several ports "for fun", so no one has tried to upstream anything. But if you have a real use case that could benefit from this, we should talk with the LLVM people and see. Feel free to file an issue and cc me and we'll find the right people.
About the containers - if you are not running them in privileged mode, you should be pretty secured, especially by limiting what kind of binaries the containers have.
The use of the word "realtime" for web tends to trigger a lot of systems developers who use the term to indicate that you can predict the actual real world time that something will take to compute, typically used in automotive and robotic settings. That being said, I didn't invent the word and words have multiple meanings and contexts. In this context it simply means push-delivered data that is pushed when updated. /disclaimer
I believe "realtime" in this case pertains to the synchronization of data amongst clients through websockets. This is how I've seen the term "realtime database" most commonly used.
With RabbitMQ, you ensure* that each task is only attempted by one worker at a time, and you don't have to do anything special to ensure that at the application level.
*I'm simplifying a bit, there are edge cases where e.g. you lose a worker that has already started a task but not completed it.
It is not similarly easy to do that with a lightweight hypervisor.
Agreed. I realized I inverted your hypervisor comment. Hypervisors have the compact contract that has any reasonable chance at being audited. Container security is basically a screen door.
For untrusted multitenant workloads in 2022, for arbitrary code and without a language-level sandbox, a shared-kernel workload isolation system might be malpractice. Again: you can easily rattle off the LPEs that would have broken a Linux shared-kernel scheme (of any realistic sort) over the last couple years.
Someone else dunked on you for a Joyent bug from a bunch of years ago. I didn't. Security researchers were dunking on shared-kernel isolation even back then, but it was a much harder decision back when the only alternative was expensive, memory-unsafe legacy hypervisors, and I would have had a hard time weighing guest escape vs kernel LPE back then too.
But this time and the last time we bounced off each other on this, we weren't talking about 2015; we're talking about today, when there are multiple memory-safe hypervisor options. I don't think it's an open question anymore.
I also don't understand how you can coherently argue that people should have "religion about Rust", but also put their faith in C-language OS kernels any more than they have to.
Further, I don't understand how Meltdown helps your argument at all here, since both isolation strategies are susceptible. Memory safety also doesn't protect you from control-plane SSRF vulnerabilities, but you can immediately see why "control-plane SSRF vulnerabilities mean memory-safety is overrated" is a bogus argument.