Containers do provide benefits.
Containers do provide benefits.
Containers are not a security measure. They're acceptable for some resource constraint considerations, but you should not ever use them because of security concerns. If that's something to worry about, use VMs at the very least.
This is repeatedly really commonly, but is bordering on wrong more and more every year. At the lowest level it's not even completely correct. Containers are resource + namespace isolation -- cgroups and namespaces -- isolation and kernel-assisted resource hiding/access control is absolutely a facet of good defence-in-depth. Enhanced security is not a core feature of containerization as a whole but it can absolutely help with your security posture.
For what % that satement is wrong, it's even more so because it fails to consider the container runtime. Container runtimes these days (containerd is the best one IMO) include ways to run containers at different levels of isolation (VMs and as micro kernels), so a move to containers can absolutely help security posture, because now you don't have to pull out packer/ansible/etc to build your VM, you can just take the same container you were running and change the runtime. Time to ship/launch improved security features is important, and containers can reduce that burden for teams.
Containers are a process isolation and resource control tool -- they can absolutely aid in security posture. Would you say the same about BSD Jails?
Just so there's some more value here, here are some projects that are somewhat close to cutting edge (though it's been a while) in this space that I like:
- nabla-containers.github.io/
- https://github.com/firecracker-microvm/firecracker-container...
- https://github.com/kata-containers/runtime (see https://github.com/kata-containers/documentation/blob/master...)
- https://rootlesscontaine.rs/
- https://www.kernel.org/doc/html/latest/admin-guide/cgroup-v2...
- https://starkandwayne.com/blog/the-capable-kernel-an-introdu...
The container ecosystem surfaces a lot of really cool resource limiting and isolation techniques and puts them in very close reach.
Well sure, this heavily depends on the people you're talking to and what their imaginations are like. If it's the "containers are lightweight VMs" crowd then sure -- that view is fundamentally wrong.
> Linux control groups provide no meaningful resource isolation.
So this is a pretty bold claim -- as far as I can see processes certainly get OOMkilled when they use more resources than allowed for by their cgroup.
Do you have a link you could share? Is some fundamental part of cgroups (v1? v2?) broken in some way I haven't heard of up until now such that everyone has patched around it and it does what it says on the tin but despite the code in the kernel?
Another example is a control group that is chronically out of memory. It may page vigorously, which will have external effects on other control groups through unaccounted resources like nvme controller time, memory fragmentation, global vmscan, etc.
A third popular way to abuse the resources of other containers is through networking. The way Linux handles network traffic is frankly hostile to proportional resource sharing. A control group with zero CPU quota can still easily cause external CPU time consumption.
> Like traditional containers, Firecracker microVMs offer fast start-up and shut-down and minimal overhead. Unlike traditional containers, however, they can provide an additional layer of isolation via the KVM hypervisor.
I think the security mechanism of note there is the hypervisor -- a VM technology. The fact that it is manageable via a daemon/API that was created for containers, does not mean it provides security via Linux containers, the mechanism.
A similar story can be told regarding capabilities, a mechanism that is orthogonal to containers -- all combinations of with and without capabilities and containers have their uses.
I do think GP is overly glib and assumes a particular threat model in order to dismiss Linux containers so thoroughly, and you are right to note "they can absolutely aid in security posture", etc. But GP is correct to note that VMs (though also imperfect) provide certain kinds of security/isolation that Linux containers, the mechanism, cannot yet match.
If you can easily break out of any containerized enviornment, I'd suggest you could register for bug bounty programmes and make quite a lot of money also you can get a reward from Jessie Frazelle by escaping from https://contained.af/
Obviously containers have a larger attack surface than say a VM hypervisor, but security is not an absolute and both containers and VMs have suffered from breakout issues in the past. No security measure in isolation provides perfect protection.
Restated, my point here is that while containers can be more complex and have pitfalls (a lot of which have been worked out somewhat at this point), there is no complexity free lunch -- `[docker|podman|crictl] run --rm --cpus 2 --memory 500mb ...` is pretty darn easy, and more so than writing properly portable and well-considered systemd unit files (and putting them in the right place, with the right permissions, under the right slice, etc). It's easier than most of the options out there (including the old methods of per-user resource segregation).
--cpus 2 --memory 500mb
to the argument vector liberates these options' definer from "understanding [...] the subsystems that power it", and reflect on their potential impact. And I'd argue that THIS is the real complexity incurred by these kinds of resource constraints that seem so simple on the surface - not the specific syntax or location you have to use to introduce them. All of which makes systemd unit files and their settings' implications (which are amazingly well-documented btw) as good as any other option, imho.If I want to run a useful piece of software like redis let's say, but I want to run it with a resource constraint to make sure that it doesn't take more than 2CPUs and 500MB of memory, it is far easier to do that with the following command line:
docker run --rm redis --cpus 2 --memory 500mb -p 6379:6379
Than to write the equivalent systemd unit file, set up the isolated filesystems that docker would let you easily bind mount in, etc. This is like comparing systemd-nspawn to systemd -- if systemd-nspawn isn't simpler than systemd then what are we even doing.Docker won because of it's developer ergonomics (containers weren't new), systemd won because of it's feature set, convenience and sturdiness. They're different tools with different primary use-cases.
As a (former) developer I d rather ran redis in a container during development, but for production I d rather rely on boring VMs unless some scaling is required. (Managed k8s case set apart)
Agree, but my view on this is that the implicit answer of how you do all that (i.e. the file system, syslog) is now gone. There will be pain (complexity) in the short term, but at the end of the day, we're going to be able to build much better orchestration and systems. To kind of restate that, before you had to worry where a process wrote out it's output (stdout? /var/log/<program>? /etc/<program>/logs? /home/<user>/<program>/logs? syslog?), now you know want to get the non-stdout/stderr logs of the thing you're running, you'd better give it a volume to write to (which may be fake, and actually write everything to some remote storage or something), and I think that's a step forward.
Of course, I'm not saying containers should go everywhere -- relying on boring VMs over containers is fine too -- but I think rich world of functionality available to container-driven workflows is popular for good and bad reasons, and the good reasons are worth exploring/beneficial to me.
I thought the kernel has virtual memory and if you consume more it will swap some memory and thats it.
Couldnt you just manage memory from inside your app? if its redis, then check the db size and shutdown/cleanup gracefully, and not crash redis with OOM?
> I thought the kernel has virtual memory and if you consume more it will swap some memory and thats it.
Well just to make sure the right memory goes to the right places -- if someone uploads a large file and you've made a mistake in your code that tries to hold it all in memory instead of buffering it straight to disk for example, you'd want that process to crash, and not your machine.
Also you generally don't want to swap, so much so that Kubernetes disables it immediately[0]. Not that Google is the only group with the right answer but they seem to think nothing can come of a machine having to swap. Maybe they're right. Even if they're not, A world where one service swaps[1] (I've never done this with docker to try it though) is probably better than one where it uses all the memory and everything swaps.
> Couldnt you just manage memory from inside your app? if its redis, then check the db size and shutdown/cleanup gracefully, and not crash redis with OOM?
You'd be surprised -- some languages just don't have a way to very easily get feedback from GC[2]. It's also something that I don't think most people think about, messing with the -XmXx<setting>s in Java is definitely year 2/3/4 java development for most people.
[0]: https://github.com/kubernetes/kubernetes/issues/53533
[1]: https://docs.docker.com/config/containers/resource_constrain...
contrast it to Ops mentality of software is a cattle (here is your memory quota and if its OOM, just kill/restart the service and hope next run it wont run OOM)
Wow, given how expensive RAM is it is no surprise they will never allow kubernetes work with swap, becausr it directly translates to $$$ for GCP and other cloud providers. Also since everyone loves using Java/Spring/Node and other memory hungry frameworks - it print enormous $ for cloud providers to require users overallocate RAM and disable swap.
or am I just spitballing conspiracy theory here and there is no conflict btw decisions like these and vendors' revenue streams?
To install:
apt install mypackage
To edit the systemd unit (automatically creating the file in the right place): systemctl edit mypackage
Then add these lines to limit to 2 cpus and 500mb memory: [Service]
CPUAccounting=true
CPUQuota=200%
MemoryAccounting=true
MemoryHigh=500M
And then to start it now and on boot: systemctl enable --now mypackage
This is now integrated with package updates, starts on boot, logs to the same place that most other system utilities do and so on. --cpus 2 --memory 500M
is easier, and gets you the same results though they may not be as permanent or as well managed -- the management and external stuff is an orthogonal concern, and that's not the situation I was addressing. The original point was insinuating that throwing up a binary and getting it to run. your filesystem is also not available to the container by default, and in this way docker sort of fails closed. If you're running a rootless container, the story is even better.One thing you have not covered is filesystem isolation, which docker also does very easily. There is a lot to configure on the systemd side[0] and the parts that are overlapping are just easier to configure and run with docker. Systemd is the better tool to build repeatable installs for pet processes, but again, there is a lot of knowledge underneath that is related. People to this day still complain that systemd does too much (I personally like it a lot, and it's great to have everything in one place).
[EDIT] Just to make myself clear, systemd is an amazing tool -- I like it, I run it, I'm not smart enough to administer a more complicated setup -- but docker is easier, for a large part of the small subset of systemd's capabilities that docker covers.
Yes, you'd need to learn how to make a systemd unit file but honestly, that shouldn't be a problem at all. I've had much worse headaches fighting Docker's networks and firewall-overruling network configuration in the past. Every time I forget to specify a restart option, I have to dig through my shell history again to see how I launched the image so I can kill and recreate it, or I have to look up that oneliner that shows you the docker run command for a running container.
I run stuff in Docker for one simple reason: I have it installed and I'm too lazy to think about software sometimes. If you're packaging software, that laziness isn't a reason to pick a distribution tool.
Static binaries have other problems, perhaps most importantly package management. How do you distribute your app? Traditional deb/rpm packages? curl-to-bash installers? Snap? Flatpak? Binaries that people flog into /usr/bin? What about automated updates, do you set up a repository, do you include a self-updater in your code? The list goes on. Docker solves that by having one general source of packages with the option of adding a repo of your own (take note, Canonical, your shitty Snap Store is practically useless without that last bit).
For services I run on servers, I much prefer traditional packages with dynamically linked executables, for a simple reason: automatically fixing security issues and bugs across applications with a single update command, instead of having to wait for every maintainer of their statically-linked tool to update their dependencies.
It's one of the big problems I have with Rust; many dependencies and applications embed their own versions of dependencies with sometimes very specific versions, and when there will eventually be a massive security problem in one of the TLS packages I'm dependent on recompiles and code modifications from random open source maintainers.
I can only offer one explanation: marketing.
Our images are made from Go binaries, so thanks to Go modules, we can manage dependencies centrally from the main module and rebuild.
I let you research when that was, and wonder what container technologies might have used since then.
I know when they make sense and when they happen to be fashion.
Half the "static" binaries that used to float around weren't even fully static because of glibc, until building with musl (and getting bit by getaddrinfo) became widespread -- the situation is not as simple as you made it seem.
My point is that I see containers being overused nowadays, like "big data" that fits on a USB pen kind of scenario.
I think we've basically stumbled into a very effective and widespread packaging paradigm though -- now you don't even have to pick the right language/toolset to get static binaries easily -- just throw a container over the wall and your filesystem requirements (mounts), network requirements, etc will be made pretty obvious to the person doing the deployment
Then I remember Linus point of view about monolitic kernels and how they are supposed to beat micro-kernels, what for, when people put hundreds of virtualization layers on top.