A Linux sysadmin's introduction to cgroups
redhat.com
redhat.com
A personal note here too. cgroups were invented at Google in the early 2000's. If you're using search today, gmail, docs, maps, etc. You're using cgroups. It sounds simple but this tech really powered a whole wave of innovation/startups/projects that almost everyone interacts with on a daily basis. Either through touching a Google service or interacting with anyone using Docker or Kubernetes (running on cgroups/namespaces). Pretty impressive.
#14 - Introduction to Linux Control Groups (Cgroups)
https://sysadmincasts.com/episodes/14-introduction-to-linux-...
#24 - Introduction to Containers on Linux using LXC
https://sysadmincasts.com/episodes/24-introduction-to-contai...
Thank you.
This tech has completely changed the sysadmin landscape/job descriptions and sort of threw tons of gas the whole devops movement.
Disclaimer: I worked at both Docker & Google. Although not on this tech specifically. Opinions are my own here.
I'm happy to find that I agree with you on at least one area of this new ecosystem :)
The reason we need an orchestration system for our orchestration system, is the orchestration system is a snowflake. We need an orchestration system because the container system is a snowflake. We need a container system because containers are a snowflake. None of it is really standard or easy to implement, because none of it was designed in the spirit of Unix. It was all just some random companies who threw some crap together so they could start making money selling ads or hosting services. (I don't mean cgroups, I mean the tools that use them, although cgroups are kind of warty in themselves)
All most people need to do is run a couple processes on a couple hosts, send packets between services (and networks), and have something start/stop/restart those processes. You can have all of that with, say, systemd (ugh), or three different standard interfaces that any program can use with standard libraries/ABIs. Notice I didn't say "with 5 different daemons running on 3 hosts that need complex configuration, constant maintenance, and a migration effort every 3 months".
The Open Container Initiative seems close to getting the first thing done. If we get the other two standardized, we can lose a whole bunch of the onion layers. Consequently a lot of people will need to find a new way to make money, because there'll be not nearly as much need to pay people to deal with the onion.
I'm not super familiar with plan 9, but given what I do know, I'm pretty sure that could be done with an rc script. I'm pretty sure that a) you can fork processes after starting them on remote systems, and b) you can query remote systems to figure out which ones have free resources. Given that, it's just a shell / rc script away. Then you just need another script to do the management. ;)
It reminds me of how electric cars are a 190 year old technology whose actual peak popularity was 120 years ago. Yet here we are, with manufacturers predicting they'll finally be popular again, in 10 to 15 years.
Still delivering on VMs, might be out of fashion, but it works, everyone gets it, and I get back home on time.
Why ugh? I do have all of that with systemd, and it’s great. Rather than running each service in a separate filesystem and requiring them to communicate over a virtual network, I use ProtectSystem and ProtectHome to keep services as isolated as they need to be. Rather than creating a network bridge and doing NAT inside my server, if services require network isolation, I use NetworkNamespacePath to assign them to namespaces that contain different VLAN interfaces, which are configured in my router just like physical interfaces. Rather than building a huge container image for every application, I use my OS’s package manager to install apps and manage their dependencies. The skills required to do this can be applied to any systemd-based Linux distribution, and are needed to properly understand and troubleshoot systemd-based Linux distributions running in containers anyway. I’d need some new tools (eg. Ansible) to scale this setup past one or two hosts, but I’m baffled by the people setting up complex Docker/Kubernetes systems just to self-host a few web apps.
Red Hat is trying to base all their stuff on Podman, which is Docker rewritten to by a script instead of a client/server (but can be renamed and used as a drop-in Docker replacement with the same command API). If they have their way, that statement will not longer be quite as correct.
Redhat’s answer to capital D Docker is cri-o.
I'm not sure that's true - a lot of Kubernetes system use containerd these days.
Why use something higher level than you need that wraps containerd (namely Docker) instead of something that is exactly at the level you want: containerd itself? The only real advantage of wrapping docker is that doing so ensures the docker cli works for introspecting containers.
The Docker folks don't want people to be wrapping dockerd, because then those infrastructure folk complain about changes/additional features in dockerd. Even dockerd is (glacially) moving in the direction of move heavily using containerd, for example using it to pull/store images.
The longer-term intent appears to be for dockerd being more for implementing docker specific stuff like its networking, swarm/stack, and various features needed by the front-end.
In my understanding, most of the initial commits (at least for LXC, which was the initial userspace) came from IBM, who funded it with interest in kernel-level resource balkanization for their largest mainframes. Google's kubernetes only appeared post-facto after docker, which itself was basically an lxc wrapper for a long time.
Source: I corresponded with two of the authors ~2009, was an early lxc user, and provided some security fixes for early docker.
Basically saying that 'containers on Linux' are 'made at Google' is not true.
Edit: I started questioning my memory of this, so I poked around some. Poul-Henning Kamp did the initial work in April, 1999, here is the commit:
https://svnweb.freebsd.org/base?view=revision&revision=46155
More details from him about it:
https://en.wikipedia.org/wiki/CP/CMS#Overview
that's exactly what jails/zones do (System, NOT Hardware virtualization)
>"Nov 1999: Alexander Tormasov visited Singapore and proposed a new direction to Sergey Beloussov: container virtualization. He formulated three main components: containers as a set of processes with namespace isolation, file system to share code/ram and isolation in resources. Indeed it was 1999 when our engineers started adding bits and pieces of containers technology to Linux kernel 2.2. Well, not exactly "containers", but rather "virtual environments" at that time -- as it often happens with new technologies, the terminology was different (the term "container" was coined by Sun only five years later, in 2004)."[1]
There were many efforts to get some kind of container technology into Linux, from the early 2000s (and probably earlier). VServer, OpenVZ, etc, were all trying to support full server virtualization.
Cgroups was different from most of them in that it wasn't trying to present a virtual server environment, just facilitate resource control. At Google, all the jobs we ran on Borg (the internal system that later inspired Kubernetes) knew that they were in a shared environment, so there was much less need to fake a private virtual environment. The low-level libraries linked into all Google binaries did a lot of coordination with Borg to make the sharing fairly painless. But we were trying to run tens of jobs on each of our increasingly-larger machines, and some of the batch jobs would inevitably end up hogging resources, which would hurt the important latency-sensitive jobs. CPU and memory isolation were the main things that we cared about, with network/disk as secondary elements.
For a while (in 2005-2007) in Borg we approximated this by using the kernel's "fake NUMA" support (originally intended for testing the NUMA code) to break up the system into a bunch of fake NUMA nodes, and using cpusets to reserve CPUs and memory chunks to important jobs. It was pretty ugly and rather coarse-grained, but it worked rather well and was the first widespread (running on millions of Google Linux servers) userspace system for controlling a cgroups-like system.
Essentially, cgroups piggy-backed on to the existing cpusets mechanism/API which had already been accepted into the kernel, and expanded it into a more generic way of creating a hierarchy of groups and mapping a process to a group. This was a lot simpler to get accepted than an entire virtualization system. Given that mapping, making other resource scheduling be group-aware rather than just process-aware was much more straightforward. The same userspace support in Borg that had been used for controlling cpusets worked pretty much as-is with cgroups, since the basic API was the same (just more resources were supported).
And as noted, OpenVZ and Linux Vserver were earlier container subsystems for Linux.
I picked a file from the patch index mail linked above, mm_inline.h, and went scrolling through its history at https://github.com/torvalds/linux/commits/master/include/lin... but didn't see a corresponding change there. I guess the patches might have gotten refactored before merging or something too but would be nice to have a pointer in the Linux tree history that would work as a reference.
edit: tried looking for linux/container.h too but that just brings up a newer acpi container related container.h: https://github.com/torvalds/linux/commits/master/include/lin...
edit 2: history for linux/cgroup.c starts in a year later in 2007, with this commit: https://github.com/torvalds/linux/commit/ddbcc7e8e50aefe467c... - it has mentions of people from OpenVZ, IBM and a bunch of unaffiliated domains in addition to the signed-off-by line from a Googler (Paul Menage).
It basically says:
Based originally on the cpuset system, extracted by Paul Menage
* Copyright (C) 2006 Google, Inc
* Copyright notices from the original cpuset code:
* --------------------------------------------------
* Copyright (C) 2003 BULL SA.
* Copyright (C) 2004-2006 Silicon Graphics, Inc.
*
* Portions derived from Patrick Mochel's sysfs code.
* sysfs is Copyright (c) 2001-3 Patrick Mochel
*
* 2003-10-10 Written by Simon Derr.
* 2003-10-22 Updates by Stephen Hemminger.
* 2004 May-July Rework by Paul Jackson.
just wow how old this is."Adding Generic Process Containers to the Linux Kernel" https://www.kernel.org/doc/ols/2007/ols2007v2-pages-45-58.pd...
Although the paper was written by Paul B. Menage it looks like Balbir Singh and Srivatsa Vaddagiri from IBM also made contributions to it.
Yes, Kubernetes uses cgroup too.
Systemd units are also based on cgroups.
Edit: well, at least part four, finally, properly document how to cope with systemd. And guess what, it is different than 90% of documentation on cgroups on the internet.
https://www.schutzwerk.com/en/43/posts/linux_container_intro...
I also tend to think that processes in cgroups are a sweet spot of lightweight containerization that can do quite a bit.
"it's not just badly written kernel/cgroup.c - the interfaces on both sides (userland and the rest of kernel) are seriously misdesigned. As far as I'm concerned, configuring it out solves my problem nicely."
That was in 2011, so things might have improved. What remains however is that cgroups was added to the kernel, by Googlers, for easier maintenance, but with an implicit understanding that no sane person would actually make use of it to do something important.
... enter SystemD.
Cpusets, which had already been accepted into the kernel a year or two previously, provided the basic userspace API for cgroups. It was just expanded to support control files for more resource types. (And multiple resource hierarchies, although that's gone in cgroups v2 I think).