Proxmox Docker Containers Monster – 13000 containers on a single host
virtualizationhowto.com
virtualizationhowto.com
And just like that you have a cluster to run containers on. I really like the simplicity of Docker Swarm. I've been using it for at least five years and it's just worked.
During the COVID lockdown I got tired of having to open a UI (at the time I was using CapRover[0]) to edit any of the services I run so I decided to make my own PaaS with a nice CLI. Connecting to the docker socket is easy and the API is simple enough. It's been working no problem for the last two years.
The only complaint I have is that I can't see the user's IP for HTTP requests[1] but there is some hope in the form of Proxy Protocol[2]. I have no idea how complex the code for Docker Swarm ingress is, but I may spend a weekend in the near future scouting the code to get an idea. The current possible solution is to put a load balancer in front of the cluster that either sets the X-Forwarded-For header (or any of the equivalent ones) or speaks Proxy Protocol but I will avoid that solution for now.
I recommend Docker Swarm as a solution for anyone starting that doesn't want to spend hours and hours configuring a production environment. Even if it is just one node, you get services, replication, healthchecks, restart policies, secrets... And it all starts with that simple command, no further config needed.
[0]: https://caprover.com/ [1]: https://github.com/moby/moby/issues/25526 [2]: https://github.com/moby/moby/issues/39465
Why Swarm over k3s?
k3s, as simple as it is compared to full k8s, still carries some of the complexity of k8s. Swarm just feels simpler and easier to manage.
Making Swarm production ready is also a time sink. Most people end up with a bunch of bash scripts mashing together YAML for secrets, deployments, etc.
Personally I think it's a waste of developer time. Most people don't need the features of Swarm or K8S. Just use a VM at a popular IaaS provider and stick a LB in front of it. Then you don't have to ask questions about IP rate limiting, or other gotchas. If you outgrow a couple of VMs, we can talk about throwing ASGs or containers at the problem. Honestly though most people I talk to don't need containerized solutions.
For version 2 they've replaced it entirely with K8s however.
I was expecting Proxmox's LXC capabilities to be used to scale up to 13000, but this is just VMs + docker allowing that. Seems like the same thing could be done with any LVM hypervisor and VMs? Can someone correct me if I'm missing something?
Comparable to building something with make -j64.
(EDIT: On Windows that's also an option, or you set it to run against Windows' native container support - which then can only run windows-based images. But really, usually people mean "on Linux" when they discuss how Docker works)
Macs don't run any containers natively, because macOS doesn't support containers and doesn't have a Linux ABI translation layer (some operating systems, like Illumos and NetBSD, do, and you can run Linux binaries on them almost like you can run Windows binaries on Linux via WINE). On M1, when you run aarch64 containers under Docker, you are 'merely' emulating an environment for a Linux kernel to run in and you can pass certain aspects of your CPU through. But when you do the same with x86_64 containers, you are additionally 'emulating' (via a translation layer) x86_64.
Hopefully some day macOS will have its own truly native containers, like Linux and Windows do. That could be a serious savings for containerized macOS development environments, because they will be way more efficient than Docker Desktop on Mac currently can be. For now, the only way to get that kind of efficiency is Nix— which is great and better suited to that use case than Docker, but underutilized at many organizations.
Probably even then people will still use virtualized Linux, since that's a closer environment to where apps will usually run in prod. But at such an org I'd argue for using macOS containers for local development, should they ever come to be available. The efficiency gains are well worth it IME.
Because it's easier to use a CPU emulator on just one kernel a single VM than manage multiple VMs. So you can use QEMU just like Macs already use Rosetta for x86_64 macOS binaries on aarch64 macOS.
Linux further has a built-in feature for running 'foreign' (different CPU architecture, but still Linux) binaries (and more!) in a transparent way, so if you set this up in the VM you can invoke those x86_64 binaries as if they were native, without really thinking about it: https://www.kernel.org/doc/html/latest/admin-guide/binfmt-mi...
You can also use it with DOS and Windows binaries, or via Apple's Rosetta for Linux instead of QEMU for CPU emulation (and there are some other open-source competitors in that space as well). Pretty neat!
There's a Linux VM hosted on your mac. There's a docker daemon running on your mac. The docker command process above communicates with said daemon which has the ability to exec processes inside the Linux VM. It uses that capability to exec runc (or something similar) which in turn starts a bash process inside a namespace inside the VM.
If you run some other container the same thing happens, using the same VM.
So the secret is : docker isn't running anything on your mac, it's running stuff on a Linux VM hosted on your mac.
Needless to say, this doesn't make any sense on macOS because there's no Linux kernel. Therefore you need a Linux VM to run Linux containers on macOS.
When it comes to running for example an Ubuntu docker container, it would share the Linux kernel running in the VM with other containers on the same host. The "Ubuntu" is the distribution packaged specifically for containers, meaning that the regular Ubuntu packaged kernel is not used and neither is the init system because the Linux distribution container images are meant to containerize single applications and traditionally do not need the concept of "services".
I'm sure it's possible to run systemd inside a container, just as some people run Cron and even X11 inside containers.
Contrast this to 'operating system containers' which are designed to run an init system, so individual containers are intentionally multiprocess. Many virtual private servers run in containers of this kind, e.g., by OpenVZ.
On that note, how come you've gone with Docker for this rather than something like LXC or OpenVZ, which are ostensibly designed for the large, multiprocess container use case?
I would be very keen to understand why LXC is better than Docker for this. I've also read this kind of "marketing" around so called "operating system containers" but I am still clueless why they are better. Concretely, what are the benefits?
And on FreeBSD there are Jails and BHyve, so not a Proxmox audience as well, IMHO.
It comes available with an install of Docker, is easy to setup and operate, has great optional UI solutions like Portainer (analogue to Rancher for Kubernetes), has one of the lower resource usages for the orchestrator itself, as well as supports the Docker Compose specification, which in my opinion is far more usable than the Kubernetes manifests (though less powerful than Helm charts) and far more common than Nomad's HCL.
For my Master's Degree, I explored a comparison where I ran the same workloads across a Docker Swarm cluster and a K3s cluster (a great Kubernetes distro that's low on resource usage as well) and even then Swarm used less memory (~2x less than Kubernetes for the leader node both under load and when idle) and used a bit less CPU (~30% less for the leader nodes under load) as well. That said, K3s still performed admirably, at least in comparison to RKE which wouldn't even run in a stable fashion on the limited hardware that I had at the time.
Maybe one of these days I should run Proxmox in my homelab as well, instead of just something like Debian or Ubuntu directly on the hardware. Also, while Podman is great, Docker still seems like a dependable option just because of how common it is and given how it's gotten more stable over time (despite the arguable architecture disadvantages).
I think the only actual issues I've had since when using Docker Swarm have been using a network that ran out of addresses to assign to the containers (probably some default), some Oracle Linux bug where kswapd would top out the CPU when the swap got full, as well as some Debian bug years ago on an old version of Docker that caused networking to fail and the cluster needed to be re-created to fix it.
However, though swarm is easy to use and configure, I've found that having a small swarm (e.g. three nodes) means that you end up having each node be a manager as the swarm fails if half the manager nodes aren't up.
I've also found that "docker compose config" doesn't work in a compatible manner and have to use "docker-compose config" instead.
have you found why it's so? I'm curious - in theory all the processes inside are the same, except, what it should be slight overhead, from Swarm/K3S itself.
What I saw during benchmarking was that the load on the worker nodes was almost the same for both K3s and Swarm, presumably because most of the resources were used by the actual container runtimes and containers themselves, the overhead for communication with the orchestrator not being too notable.
However, when I tested the whole setup under load (containers serving lots of web requests, databases inside of containers etc., some containers eventually getting OOM killed and restarted), I found that the leader nodes had more pronounced resource usage differences, as described above.
My guess is that Kubernetes (even K3s) is just a bigger and more complex system than Swarm, which probably also explains the comparatively higher requirements/overhead (for even single node Kubernetes as well). That's also why I'd be careful about running anything apart from the orchestrator on the leader nodes.
Not knocking this achievement though, it's awesome they were able to pack that many containers in one host.
LXC is just a container, docker is much more than just that. It is a recipe (dockerfiles), it is sort of a social project sharing setup, it is like a version control system (the layered filesystem that only updates changes), it is a disposable dev container, it is a deployable runtime container, etc
I haven't used TrueNAS since it was FreeNAS and BSD based.
LXC is a first class citizen in Proxmox, with clear documentation, config files, CLI tools and web UI. It can be backed-up just like a VM (Proxmox Backup Server FTW!), can be set up mostly like a VM, it boots instantly and so on. It's such a joy to work with, unless you want access to actual hardware (which includes things like mounting or hosting NFS/Samba), but even then it's easy to find help on docs and forums (the latter are surprisingly up to date).
But LXC does not have that "ephemeral" nature like Docker. Container "templates" are clean but full operating systems, full hard drives are connected and host all data, both user files/images and OS. Just like a VM, I guess.
Now, you can actually run Docker inside LXC easily and have best of both worlds: docker with docker-compose and alike to quickly prototype and homelab AND super light and quick to boot machines which will host that Docker for you.
It's actually a very clean and pleasant approach, I highly recommend this as tool for testing and homelabing.
Is it for making a RAID directly on the HW drives?
For me so far there is always a decent workaround when I want to use USB device etc.. The only thing I am currently stuck with is sound in LXC container (I build kind of a VDI desktop, RDP access to a system with KDE) and struggle to pass the actual sound card through. There is some active proxmox forum thread on that so I'm hopeful.
And yes as others have said, LXC and docker use cases seem different.
LXC seems to be for the same use case as VMs, but smaller, easier to spin up, better resource utilization, etc.
Docker can be used that way too, but is more useful with the surrounding ecosystem of being ephemeral (all data stored on remote storage, container only contains the "app") and with supporting deployment infrastructure.
edit: Wait, did they install docker/portainer on Proxmox bare metal? They say to access Portainer through the Proxmox host IP, but any CT or VM created on Proxmox would probably have its own IP on a Proxmox bridge. So the IP should be the IP of the VM/CT hosting the docker install, not the Proxmox host.
That said, containers are very lean (or can be with the right setup) given there is no kernel, drivers, etc to load.
[1]: https://www.ongres.com/blog/63-node-eks-cluster-running-on-a...
When I first saw Proxmox I also wanted to see if I could use it to manage docker containers, but it doesn’t support it directly. For working with containers you need other tools, eg Kubernetes.