Build your own Docker with Linux namespaces, cgroups, and chroot
akashrajpurohit.com
akashrajpurohit.com
Why is it still necessary to have whole, full-blown OS filesystems inside of our containers, if their purpose is running a single binary?
Dependencies/dynamic libraries are decent reason, sure. But wouldn't it make more sense to do things "bottom-up"? i.e. starting from an empty filesystem, and then progressively adding the files that are absolutely necessary for the binary to work, instead of the "top-down" approach, which starts from a complete OS filesystem and then starts removing the things that are not needed?
Because the second you want to do anything involving https, you need certificates, and that's where having a minimal but existent base image starts shining, and mostly goes up from there...
If it's one single golang binary, or a weird python container running only on specific version of Debian compiled during a blue moon, or some 8GB java enterprise bloatware, it's the same.
"or a weird python container running only on specific version of Debian compiled during a blue moon"
This.
I guess I've been fortunate that I'm able to reject software like this from my stack. I know that everyone isn't so lucky.
I mostly do nodejs, and have zero need for containers when a simple npm install gets all deps.
Or if I need performance, a single go binary.
I tried doing the container thing just to understand how it all works and what the hype is about.
It seemed needlessly complex and hard to develop/debug.
For more complex situations where you need a bunch of interacting programs and services, I prefer stuff like Ansible and VMs, or just manually setting up a base image.
I guess it's a "get off my lawn" kind of thing. It seems like containers are used a lot by folks who don't want to learn ops, like how to install and configure postgres, redis, etc...
I think that's a mistake, and just pushes the problem onto others who have to support the software in production.
docker build -t my-app . && docker push my-app
Then all of a sudden it's a reproducible, reusable deployment script that works for any language, any application, and that any other dev on your team can run (as opposed to playing "which flavor of coreutils did they use when they wrote this?"). It's provider agnostic - you can run it on DO, AWS, GCP, whatver. You get free rolling/blue green/canary/whatever you prefer deployments, a "basic" cross compilation out of the box. They're not magic, but they are an excellent abstraction, despite the warts.There doesn't seem to be a _huge_ drawback to running an Ubuntu image vs Alpine, aside from the image size.
I switched from using Alpine for my own images to using Ubuntu instead a while ago: https://blog.kronis.dev/articles/using-ubuntu-as-the-base-fo...
So far, it's surprisingly nice. Ubuntu based images take up more space, but thanks to layer reuse and storage being affordable, this isn't a big issue in practice (at my scale).
There are no surprises in regards to performance or package management that I have to deal with. Actually, I can use Ubuntu/Linux Mint locally and have the same install instructions for tools/dependencies as well (if I want to test things outside of containers, running on the system directly).
It's also really nice to be able to build my own base image with whatever tools I want available (e.g. nano and some debugging stuff) and have them be available in all of the language/stack-specific images that I build later.
The EOL is also pretty long and I don't think that there's any shame in having a "good enough" and somewhat boring distro for most dev stuff, only occasionally looking at PPAs for newer stuff.
That said, Alpine and Debian are both fine as well! I still run Bitnami images for most complex software like databases, which I think used Debian as a base: https://bitnami.com/stacks/containers
Making distroless containers is the way to go.
A RHEL 9 universal base image, which you can add just the packages you need to, is 217MB.
And if you have hundreds of images on your system, maybe remove some you aren't using. I doubt you're using them all, and if you are using them all, they likely aren't using all the space you think they are because they're probably sharing layers.
I'm not saying to necessarily use it, but I wanted to pick something that was on the heavier end of the spectrum to forestall complaints about how for some people's workloads it's not realistic to run a 5MB alpine image given their work environment and needs.
On top of this, as some sibling commenters were saying, it is a waste of bandwidth, time and disk space.
If you aren't shipping setuid binaries in your container, even if someone gets a full shell in the container they are locked down by the container permissions and cgroup limitations, which is the whole point. That some extra utility or library exists but is not running that has an exploit is really not providing much of an attack vector. If an attacker can arbitrarily run something the fact that some code is on the system is not likely to really change their abilities, unless it's setuid, so get rid of that stuff.
As per the dependencies of important projects: yep, I read their release notes. And since I do no like to do it, I try to keep them to a minimum, unless it is a prototype or a research project.
I get everyone has his sensibility, but please next time try to respect those who have a different attitude. Nobody attacked you or wants to.
Having an OS as the container doesn't mean shipping with everything. bash and awk are commnon, and curl might be for some as well since it may be a dependency for something else, buy jq won't be, as well as many other things (and often you have to include the few things I noted anyway, as many applications will call external items and strict exec use isn't always the norm).
> I get everyone has his sensibility, but please next time try to respect those who have a different attitude. Nobody attacked you or wants to.
Please consider that perhaps you're being a bit too sensitive, given this is a text medium and you can't hear my inflection.
I understand I may have sounded a bit facetious, but I was being honest there. It's a lot of work keeping up with dependencies, so good luck with that if you, I wouldn't want to do it myself, and if you are doing that, I'm always interested in hearing about ways in which people manage it[1]. It's directly relevant to my job. If you weren't doing that, then it was meant as a gentle note that hey, you really should be if you're concerned about security, because anyone that's statically compiling external requirements into a single binary and isn't has got much more to worry about then whether there's other utilities shipped in their container.
1: For example, I wrote https://news.ycombinator.com/item?id=36450815 just the other day, which is directly relevant to that idea.
If your program doesn't do that - e.g. most Rust and Go programs which are a single statically linked binary - then there's very little reason to use Docker in the first place.
People will say "but you still want to containerise things!" or "what about orchestration?" and sure those are some incidental benefits now, but they only really happened because everyone was more or less forced to use Docker to actually run apps.
Docker is a workaround for shitty software.
This just needs to be repeated for the pythonistas sitting in the back.
"Docker is static linking for millennials"
:D
As an aside, the approach you described is how Nix works, and you can use Nix to produce OCI images.
Why would you want a shell in a container? Surely the idea of a container is to run a program, not act as interactive environment.
The shell can also be useful for certain scripting steps during image build, not all containers are just going to copy in a statically compiled binary etc. In fact, I'd say the latter is significantly rarer than the former, at least in my experience.
Being able to do commands like:
$ docker exec -it mycontainer sh
is generally very helpful during development.
> https://docs.docker.com/engine/reference/commandline/exec/
FROM Scratch - starts with a totally blank image, this is the smallest option but your application must work with no supporting files (e.g. statically compiled binaries).
Distroless - has a small number of standard OS support files but no package manager, so works where you don't need to install many OS packages.
Wolfi - Newer than the others, they're building an ecosystem of minimal images for specific purposes.
* pip is not guaranteed to get you to a working state even if you run the same commands, so don't even think about it
And more. In my experience, depending on the quality of the packages you happen to be consuming, you may end up with a container twice as large as a comparable (say, Alpine) container. For example, I once tried to bring git into my Nix container and was surprised to see over 300MB in increased size.
After having built Nix containers exclusively for a year now, I wouldn't really recommend it to others unless you're willing to invest a lot of additional time cleaning up community packages.
Another great thing is that you can can easily modify existing package to remove some of their dependencies that you don't use.
This is shown here with packaging redis into a container here and making it 3x smaller than the alpine version: https://nixos.org/#asciinema-demo-cover
There's of course possibility that someone specified build dependency as a runtime dependency, but if that was done that's a bug.
If you know what you are doing you have several alternatives:
1. Just link your binary statically.
2. Set the path to the picked out special libs as an environment variable only for the binary.
3. Re-think what's wrong with your system that leads to the pain of you wanting other lib versions for that binary.
2. You can absolutely limit the system libraries that are included to only those needed. Some platforms make this easier or harder than others.
3. That you have other software that you want to run in parallel and don't want to spin up full virtualization or manage various chroot structures to organize.
I'd also add that just because a tool makes it easy to do something, and means you can bloat your file system doesn't make it inherently bad. Not everyone is trying to run a large database on a potato.
Using dockerTools from nixpkgs is much better and gives you much smaller images closer to Alpine size.
Edit: with nothing in the contents it's 144M, which is getting reasonable but still nearly 30x alpine base
It's not. This is just the path of least resistance for most container builds since many applications require various supporting files to function properly. Most of my Go containers only contain 2 files: the go binary and ca certificates.
> Dependencies/dynamic libraries are decent reason, sure. But wouldn't it make more sense to do things "bottom-up"?
In theory, yes. With dynamic loading and support for loading plugins and runtime this is virtually impossible to do. That's why starting from a base image that already includes everything you need is generally the chosen method. BSD jails/chroot operate this way and it's an absolute pain to setup for complicated applications.
I've found these after some quick googling:
https://unikraft.org/ https://hermitcore.org/ https://nanos.org/
Seems to be a living concept still, just not in the mainstream.
The only "example" I can think of off hand is Firecracker, but I'm not 100% confident that Firecracker is technically a unikernels.
My guess is that there's a number of unikernel implementations behind closed doors that see heavy use
Alternatively, instead of --directory=/ you could specify some other directory that contains an OS image (such as --directory=/var/lib/machines/debian-bookworm or --directory=/var/lib/machines/fedora-38). Multiple containers can transparently share the same image, since all the writes go to a per–container tmpfs.
https://0pointer.net/blog/running-an-container-off-the-host-...
But, if it mounts everything, wouldn't that also make container escapes very easy?
Unless you're using Go or a C/C++ stack that does static linking and thus needs no dependencies including libc, that's yak shaving to an extreme degree.
But they're really not. This is a common misconception because of the user friendliness of Docker, but underneath all that there are many moving parts that the Linux kernel exposes, which are tricky to manage manually, and additional features of Docker itself (Dockerfile, layered images, distribution, etc.). Sure, you can write a shell script that does a tiny fraction of what container tools do, but then you'd have a half-baked solution that reinvents the wheel because... it's lighterweight?
What TFA is doing is fine if you want to learn how containers work and impress your colleagues, but for real world usage, stick to the established tooling.
Like others mentioned, there are ways to do what you propose, and you can always create your own container tool based on that shell script that automates this :). Though you might be interested in unikernels, which is an extreme version of that approach using VMs.
Windows had an implementation of the Linux kernel interface for this exact purpose (WSL 1) but because of performance and compatibility challenges, they switched to plain virtual machines.
These days, virtual machine run with almost 0 overhead anyway, so if you want to run a Linux binary they're a much simpler option than implementing a whole operating system.
I switched at home FT when I saw ads in my start menu search results that first time. The Edge nags weren't enough, the pre-installed games, etc. But an ad before what I was looking for on my local system, that was too far. I realize it was only a "test" feature, but even that someone wanted to test such a thing sickens me.
I've still had to sometimes run under Windows at work. I prefer the M1/M2 macs now, just for silence and battery life, but they don't run Docker nearly as well, at least most of the x86_64 container issues are resolved, and most of what I touch has aarch64 targets.
lxc really only covers starting a process as a container with some basic configuration. Later on Docker developed libcontainer which gave it interaction with other namespacing technologies like IPC, network, etc. The interfaces with these other namespacing technologies are not the same across linuxes which is what I mean by "run anywhere".
Both the Linux and Windows kernels support containers, but the way containers work is by running on the parent kernel (as opposed to traditional/micro VM's or unikernels).
So if you want to run a Linux container on windows you need a VM running the Linux kernel to provide the host kernel for the container.
The same would be true even on Linux if you needed to use a different kernel for the container than what the parent system is running
It's not and Docker carries plenty of unnecessary complexity.
Get rid of those bloated OSes and run apps directly on HW/vHW
I understand that there are some good use cases for containers, but I still haven't encountered a need for them.
I run everything from a /home/app folder, with the app user and restricted permissions.
I avoid projects that have container-only deployment (supabase, for example).
Whenever I peek behind the curtain, I end up horrified at the complexity.
Honestly, I just can't imagine why someone would want to run postgres, for example, in a container. Seems like a nightmare for maintenance and production support.
Your alpine binary can run in an alpine container. But folks run that alpine container on ubuntu 18.04 or debian 11 or macos or wsl2.
also, docker brings a lot of usefulness to this mix - dockerfile "recipe" that builds on other recipes, layered filesystem sort of like version control, global namespace, etc
> Remember, namespaces are a powerful feature that requires careful configuration and management. With proper knowledge and implementation, you can harness the full potential of Linux namespaces to create robust and secure systems.
A clue from this post itself is that all the links were added to the intro because GPT won’t intersperse links throughout.
e:Softened my language since there’s no way to know, and w/e ChatGPT is smart anyway. Better to judge content on the merits anyway imo
Quite happy with this workflow because it helps me publish articles more frequently where I don't have to worry about stuff other than just dumping my thoughts in raw format.
Its similar to how I use Astro as a tool to generate static pages from these markdown files to easily deploy on web or TailwindCSS etc etc you get the point.
https://media.tenor.com/1PMq-CFZno4AAAAC/avengers-endgame-hu...
It's the ease of extension of the container image format that is as much responsible for the popularity of container based architectures as it is clever use of namespaces, cgroups and chroot for "robust isolation, resource management, and security". Without the image format, Docker is way less interesting and you arguably haven't "built your own Docker".
Exactly this.
Docker might have a ton of nifty features, but it's killer feature is undoubtedly app packaging and deploying. Features like chroot are already as old as time, but they never became nearly as popular as Docker for a good reason: they don't solve the problem most people need to get out of the way before being able to containerize apps.
The topic is right in my area of interest, but I don’t think want to be reading ChatGPT articles from here on out if I can avoid it.
I am exploring Linux myself for this year and hence more focused for content around that these days.
Given said that, completely understand your sentiment here, so feel free to skip it, no hard feelings but I'm gonna continue with this workflow till I find something better to improve upon this as well. :)
But I find its wishy-washy tone tiring, lacking personality or spark. I'm not entirely sure it's the AI that's bothersome to me. Reading long Wikipedia articles is not much fun either. The committee process washes all the texture from the piece.
I think what we eventually seek in writing is character and that's missing.
It's an interesting "hands on learning" style of exercise, although I'm not sure why we need a new article like this to make the rounds every so often
debootstrap focal ./ubuntu-rootfs http://archive.ubuntu.com/ubuntu/
systemd-nspawn -D ./ubuntu-rootfs
?You can have nginx, Apache, php, Perl, python, etc… siloed into the same root as your application and have multiple instances for every application simply by compiling with the “path” parameter in the configuration.
Docker have features such as layered images, networking, container orchestration, and extensive tooling that make it a powerful and versatile solution for deploying applications.
More complex sandboxing techniques include opening handles for sockets, pipes, files, etc and then hardening seccomp filters on top to prevent any new handles being opened. In this way, some containers can read/write defined files on a volume without having any ability to otherwise interact with file systems such as opening new files (all file system related system calls could be disabled).
[1] https://github.com/moby/moby/blob/master/profiles/seccomp/de...
[2] https://docs.docker.com/engine/security/#linux-kernel-capabi...
(Note that the article also doesn't to link to any specific product, just to the Docker company front page. For me the first link the front page offers is "Docker Desktop for Linux" which is I guess a virtualization based system)
I tried `unshare --uts --pid --net --mount --ipc --fork` but it failed due to permissions. `sudo unshare --uts --pid --net --mount --ipc --fork` left me in an environment where I was still in my home directory, able to see all the files and create new files which would persist after exiting.
I guess there are many other tutorials which would explain this in depth, but this blog post did not really teach me anything useful about `unshare`.
Of course, it's not feasible to add everything into a single blog, it's upon the reader's curiosity to explore more.
PS: I find man pages a lot helpful now, and would recommend the same for others as well.
[1] Namespaces: https://akashrajpurohit.com/blog/linux-namespaces-isolating-...
[2] Cgroups: https://akashrajpurohit.com/blog/linux-control-groups-finetu...
[3] Chroot: https://akashrajpurohit.com/blog/how-to-create-a-restricted-...