Life of a Container (2020)
indradhanush.github.io
indradhanush.github.io
EDIT: I found this link at the bottom of the OP's blog that was even better at simplifying things, https://jvns.ca/blog/2016/10/10/what-even-is-a-container/
"a kind of lightweight VM" seems a much more intuitive answer, with limitations.
The link is useful tho, she's a good writer.
The wrong answer usually involves a lot of hemming and hawing
It's a pretty reasonable starting point as a description compared to the meaningless rejoinder which says nothing about what the point of them is.
But assuming it's an OCI container, which will be the case if you're using any common managed container runtime and not rolling your own, it's a process in its own user namespace, mount namespace, UTS namespace, cgroup namespace, IPC namespace, and time namespace. It's assigned its own hostname and IP in the new UTS namespace. It runs in a chroot in the new mount namespace to a root filesystem assembled via overlay with a single writeable layer on top of N read-only layers shared with any other container that is launched from the same image. It gets bind mounts, kernel VFS mounts, an environment, and an entrypoint command from a config file colocated with the root filesystem in a container bundle created from the image defaults, system defaults, and command line overrides.
The simplest way to compare it to a VM is a container is the Linux kernel using namespaces and cgroups to scope the services and resources it presents to a single userspace process. A VM is a hypervisor presenting virtualized hardware and BIOS services to guest kernels.
If a zookeeper starts talking about jackdaws and crows in an interview, who’s the one being obtuse and trying to show off their knowledge?
The important thing is that the animals are taken care of and the zoo visitors are happy.
In my experience, people who need to know every little thing rarely end up knowing every little thing, and are actually the absolute worst at fixing issues of high urgency - due to needing to know every little thing.
I'd much rather have people who are able to learn fast on the fly. Those are the people who actually end up knowing the little things that are actually useful, as opposed to useless trivia.
A VM host creates a guest environment with a bunch of virtual hardware devices and starts up the guest's kernel that talks to them through its own drivers. The guest does its own hardware initialization, formats and mounts its own block storage devices, does its own bootup and process scheduling, etc.
[1] From their own FAQ: https://docs.docker.com/engine/faq/#what-does-docker-technol...
(Of course, it might not truly be a tree either, because you can use tools like nsenter to put additional processes into the PID namespace without them being children of the container's PID 1. But a tree is the common case.)
Saying it's a process is mostly correct, and saying it's a process tree is a guess even more correct I think they both miss the point.
During container setup, you get a new mount namespace so that any changes to mounted filesystems are only seen by processes sharing that new mount namespace. Then you mount the filesystem from the docker image, replacing the filesystems mounted in the root namespace.
I don't know if it'll help, but that process is very similar to what happens during boot, where you've initially got a root filesystem mounted from the initrd, but you replace it with the root filesystem from a disk after you've loaded the right drivers from the initrd.
When you run eg. an ubuntu docker image, the entire user-space of ubuntu is running inside the container. Only the kernel is shared with the host.
Traditional process groups/sessions are laughably inappropriate for this purpose since breaking out of them is trivial and in fact, about half of applications that internally use worker (sub)processes do exactly that, in order to (re)implement their own job control.
Also, you need some protection against children who do "kill(-getpid(), SIGKILL)" before the exit.
Systemd can also handle it by killing the processes by cgroup i believe, and for non services you could take advantage of that with systemd-run.
I personally think the only way to make a "detached" process should be by asking some process from an entirely different "group" (over RPC, presumably) to launch something for you: so that the newly created process would technically be a children of this other "launcher" process.
# mkdir /sys/fs/cgroup/memory/child
and
# ls -lh /sys/fs/cgroup/memory/demo/