Linux containers in 500 lines of code (2016)
blog.lizzie.io
blog.lizzie.io
It went on to waste all my tokens creating a specialized docker clone. Cool I guess.
https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)
https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)
https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)
I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.
The fact that this is possible in the first place makes me think we need a much better approach.
https://firecracker-microvm.github.io/
But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think
edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.
Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.
Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.
Kata also breaks any potential of confidential VM unless you're a virtualization wizard.
You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.
Which one is more secure? I thought gvisor but your sentence sounds like it implies the opposite.
All 3 of these are vastly different tech. Confusingly, you could run all 3 of these at the same time. I know that probally doesn't help. If you want to learn more I'd go get a linux machine somewhere like aws and play around with them.
A local agent harness is a "light" sandbox in of itself you'd also be able to learn alot from.
Bubblewrap is a good tech to check out and likely a better tool if you're looking for a simpler option. App armor in linux is also nice gvisor like tool.
Wrong assumption, my goal was to get more information from someone who sounded knowledgeable and said "it's not quite the same security level", which was ambiguous to me.
> I know that probally doesn't help.
No, it doesn't. I know how to learn myself, I don't need someone to tell me that if I want to know more, I should go read about it.
The underlying assumption behind it is that it is easier to build a secure kernel inside of golang than in C.
Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."
If a VM is not sufficient for your threat model, I'm curious what is?
Some might prefer issues that can be fixed instantly and cheap to issues that will require years and millions to be fixed.
This is the nirvana fallacy in action. Because you cannot imagine a world with perfect hardware, you think hardware that catches 99.9% of security bugs is worthless.
Like, if app in wild gets hacked, you kinda already are screwed, even if the hack is contained to the box (whether VM or container) app runs in, you still get whatever app keys app used, and you still get whatever visibility to internal network the app had. https://xkcd.com/1200/ basically.
If your app server gets hacked, all user data leaks anyway. If you divide everything to microservices so they see minimum required amount of data, attacker can still see everything that goes thru it
Also defense in depth is a thing, so even if they can be vulnerated, adding multiple layers does provide additive security (provided the holes are not frequent and cheap enough to make a swiss cheese)
#!/bin/sh
# Confine dot in bubblewrap, taking input from stdin and writing PNG
# output to stdout.
# 64 megs seems to be enough, 21 megs isn’t.
address_space=64001000
# With zero --fsize, we can’t write the output file on stdout if it's
# redirected to a file, but you can pipe it to `cat`.
file_size=0
cpu_seconds=5
# Apparently Pango or fontconfig is multithreaded now‽
# (process:2): GLib-ERROR **: 00:37:18.600: creating thread '[pango] FcInit': Error creating thread: Resource temporarily unavailable
processes=4
# We’re using --unshare-user, etc., explicitly, because --unshare-all
# uses the wimpy --unshare-user-try and --unshare-cgroup-try options.
# --remount-ro / prevents malicious code from filling the root
# filesystem with empty files.
exec bwrap \
--ro-bind /bin /bin \
--ro-bind /lib /lib \
--ro-bind /lib64 /lib64 \
--ro-bind /sbin /sbin \
--ro-bind /usr/lib /usr/lib \
--ro-bind /usr/share/fonts /usr/share/fonts \
--ro-bind /var/cache/fontconfig /var/cache/fontconfig \
--ro-bind /etc/fonts /etc/fonts \
--remount-ro / \
--unshare-user --unshare-ipc --unshare-pid --unshare-net --unshare-uts \
--unshare-cgroup --die-with-parent --new-session --cap-drop ALL \
--clearenv --setenv PATH /bin \
prlimit --as="$address_space" --fsize="$file_size" \
--cpu="$cpu_seconds" --nproc="$processes" \
dot -Tpng -Gdpi=192
# For testing, to verify that network access is indeed blocked:
# nc.traditional -v -v 127.0.0.1 8000
Still, this is enough code that I'm not sure I haven't left something out. Still pending: run ImageMagick or netpbm inside the sandbox to convert the PNG file into a PPM or BMP — there have been CVEs in libpng in the past, and of course it's potentially vulnerable to zip bombs.Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.
All those game mashups made recently by LLMs... Give me a port of Firefox on Plan 9, I'll be impressed then.
bwrap --unshare-user --ro-bind /bin /bin --ro-bind /lib /lib \
--ro-bind /lib64 /lib64 /bin/shSo, like many linux userspace applications, containers in linux are just a thin wrapper over kernel functionality.