Containers vs. pods – taking a deeper look
iximiuz.com
iximiuz.com
Not that it is an alternative to K8s in every scenario, but it should have gotten at least a mention in the intro section.
With an LXC client on my laptop I am able to launch instances without using local resources. The instance has profiles setup to add the containers to my local lan for ease of access.
Launching a new machines is only a few commands away.
This does makes me want to build "kilx" k8s inside lxd heh
https://github.com/ochinchina/supervisord
You can configure it like supervisord with a single config file,
https://riptutorial.com/docker/example/14132/dockerfile-plus...
But the golang version is much smaller / less deps and can be installed as a static binary.
If I'm writing a docker compose file, maybe I write something with a DB container, and 3x backend containers, and 2x frontend containers. If I'm using Docker Swarm, then that's probably what I deploy as a "stack."
But there's no immediate analogue of this in Kubernetes. Are my DB, backend, and frontend containers grouped in the same pod? ("Just like you can do with docker compose" as you say.) No, absolutely not! Otherwise I'm scaling up DBs when I scale up my frontend/backend. The DB is only finally grouped with them at the namespace level.
OK, so split off the DB. Do I also split the frontend/backend? Well, maybe not. Both for performance reasons that don't exist in compose/swarm and for practicality of scaling, maybe I tie those together so there's always 1:1 frontend:backend, even if I'm wasting some resources. Or maybe I don't, and put all three as separate workloads in a namespace so I can scale each appropriately.
But is a namespace like a Docker Swarm stack? Oh dear no, it's also the basis for access control (sometimes) and resource discovery (sometimes) and for service discovery (almost always, at least).
How many namespaces do I need? Where should I split stuff? Are namespaces for isolating different environments (staging, prod, feature branches)? For isolating business units? For isolating all the way down to the workload level, with only one deployment + configmap + serviceaccount + etc? To this day I don't know, even talking to the same ops folks I see their opinions change over time, and not consistently in the same direction.
So in the end there's nothing in Kubernetes that really looks like a Docker compose file / Docker swarm "stack", and that kind of sucks because that's really the "project" or "application" level a lot of developers outside microservice hellworld or FAANG hyperspecialization work with on a day-to-day basis. And there doesn't seem to be a congealing best practice over how to structure full applications in spite of this. So for anyone who just wants to orchestrate a couple dozen containers for a 10 person team, you're really just throwing YAML in a despair hole and hoping the next tool you want doesn't have a different opinion than the five you already have.
In a large organization, you might see:
- a cluster per environment (staging, dev, prod, ...)
- a namespace per team (or business unit)
- some more namespaces for operators like cert-manager, tekton, ...
- an Helm chart for each application deployed in the team's namespace
If you're building a SaaS product, you might see: - a namespace per client
- an helm chart for your product deployed in each namespace
But basically, you'll often have the following: - one "service" per container (frontend, backend, db, ...)
- one container per pod + sidecar containers like kube-proxy, or vault (to refresh secrets as they change)
- a deployment/statefulset/daemonset to scale the pods
- a service resource to access the pods (DNS resolution)
- maybe an Ingres resource to expose your services to outside trafic
Your Helm chart (or any other packaging solution) may group multiple Deployments/StatefulSets/DaemonSets and Services into a single package, describing your application. Then you deploy however you want this package.Personally, I use namespaces to split things by features:
- one namespace for monitoring (kubirds, prometheus, ...)
- one namespace for secrets (cert-manager, kubevault, ...)
- one namespace for databases (kubedb, ...)
- one namespace for cicd (tekton, kapp-controller, ...)
- etc...
Then one namespace per "project" (a set of applications and resources provided by the operators running in the "feature" namespaces).TL;DR: There is no "one true solution/pattern", only what fits your needs.
But yeah, when I said "one service per container" I was thinking about the step before deploying to K8s. It's because I've seen a few years ago some python projects completely encapsulated in a Docker container, with an init system like supervisord (running a redis, a django app, and celery workers).
> Kubernetes is a framework... There is no "one true solution/pattern", only what fits your needs.
Right, this is the opposite of what "framework" means in the rest of software development. If I choose RoR or Spring, it's because I want someone else to have made these decisions, reasonably, for me.
Kubernetes is instead like a random pile of tools, most of which are not directly useful to developers, with minimal guidance on which you should use or how you should put them together. All the sharp edges of the Unix shell, and none of the composability or fluency.
If you want security with containers, use Firecracker. It uses Micro VMs rather than just kernel-level restrictions, so even a Linux kernel security bug shouldn't be able to jump out to the host or other containers/Firecrackers.
Completely disagree. How much experience/exposure do you have to kernel security that you say is crap?
> It uses Micro VMs rather than just kernel-level restrictions, so even a Linux kernel security bug shouldn't be able to jump out to the host or other containers/Firecrackers.
You realize that firecracker uses KVM, which is part of the "crap" kernel that you don't trust? A "Linux kernel security bug" could absolutely allow a Firecracker VM to jump to the host or other containers/Firecrackers.
If your complaint is that container implementations leave the hardening scope to other tools, then sure, but I would argue that's just philosophy difference between the unix approach of do one thing and do it well, and chain tools together to solve problems, and the approach of one program to rule them all.
The hypervisor isolates guest kernel bugs from the host by nature of strictly controlling resource use from the lowest level. There are of course hypervisor bugs that allow breakouts, but they are a couple orders of magnitude rarer than the typical Linux privesc bug.
Good for you, seriously.
But that's not why people use containers. People use containers because they want to deploy random crap from the internet at the press of a button. I'd wager "rootless" is a bug, not a feature in this scenario.
No, it really isn't. At all. In fact, your comment is so out of touch it actually reads like a poor attempt at trolling.
People use containers because they offer an easy and very convenient and self-contained platform to build, deploy, configure, and run multiple instances of the same application, regardless of node or platform.
In fact, if you ever manage to get any experience with containerized applications you'll eventually notice that in all containerized apps the bulk of applications, specially microservice-based applications, are comprised of apps developed in-house.
On top of that, there is a wealth of container orchestration systems that provide support for blue-green deployments, autoscaling, system introspection, auditing, and even secrets management, not to mention networking.
Your assertion makes as much sense as claiming that people use Linux distros a because they want to deploy random crap from the internet at the press of a button, just because they provide a package management system.
> Your assertion makes as much sense as claiming that people use Linux distros a because they want to deploy random crap from the internet at the press of a button, just because they provide a package management system.
If a major reason were using Linux was for the AUR, then that'd be a valid criticism. In practice, most distros ship a package manager that only pulls from official repos, and changing that requires jumping through hoops (even Arch with the AUR requires manual work to build an AUR package). And in practice, a lot of people are pulling random unofficial images off Docker Hub (and gcr.io and such) and just running them without review.
Containers might have provided a convenient way to package, distribute, and deploy software, but that's just a nice-to-have complementing the whole reason anyone adopts containers.
To put things in perspective, while Docker is practically a household name, does snap ring a bell to anyone? Does anyone bother at all with snap?
> But to ignore the popularity of Docker Hub and claim that people aren't also jumping head first into containers because they make it easy to grab random unvetted binaries is a step to far.
This assertion is proven false with the inception of container registry services provided by service providers such as GitLab, GitHub, and all major cloud service providers such as AWS, Google Cloud, and Azure, and even lesser services such as Alibaba cloud and IBM cloud. These services might make container images available to the public, but their main role is to allow people like you and me to push their own container images to be pulled by your container orchestration services. In fact, even if you adopt third-party services you end up either using the official images made available by each project or repackaging the software yourself to reflect your own config and deployment needs.
> I'd wager "rootless" is a bug, not a feature in this scenario.
You would be mistaken. Containers don't have any magic that makes it easier or harder to run as root. In this respect, they're just Linux processes, and an administrator can run them as root or not. And like Linux processes, the widely-understood best practice is to run them without root, and indeed many orchestrators require you to explicitly opt-in to "privileged execution".
As point of fact, containers have strictly more security layers than vanilla Linux processes. They are typically thought to have weaker isolation properties than VMs, which is why we (as an industry) invariably run containers (and vanilla Linux processes) inside of VMs or forego multi-tenancy altogether.
Of course nobody vets or even looks at the mess inside the compose file, and most of this software won't even run without root privileges. (Because it hooks into various system bits and violates all sorts of isolation rules.)
People value Docker as a packaging tool; especially as a go-to tool for packaging legacy crap and software-as-a-pet systems.
Running this stuff without any sort of checking and as root is bonkers, but it is what it is.
We're kind of back in the Windows 95 era of packaging software as far as server backends go. Maybe it will change after some very serious worms and viruses his the Docker ecosystem. (Windows changed very slowly and only after tremendous pressure from cybercrime.)
I don't know about that. I recently was tasked with installing a web service, running in a container.
First thing I noticed was that the container was bundling a nginx reverse proxy, a node js server, a redis database, a rabbit-mq service and a pgsql database. I noticed because after installing the whole thing I wondered why the database was working fine even when the env variables for the pgsql were wrong. Those env variables are in the README (along with other variables holding secrets so it's not clear what is necessary and what is not from reading the documentation).
Then trying to proxying the thing I, after an afternoon of banging my head, I remembered docker has an impact on iptables rules. The documentation and support forums mixes both nginx as the reverse proxy/frontend running inside the container and nginx running in front of it.
So, documentation was not good and I don't see how anyone downloading this docker-compose.yml file would be able to run it without digging through it but it doesn't mean a well documented shipped docker bundle with 1 process per container is hard or impossible or not desirable.
I do agree that the tendency of some projects to ship a docker-compose.yml with one service running multiple processes (databases, server, proxy, cache, etc.) is a giant PITA. I mostly see that from opensource projects with a paid enterprise version. I am pretty sure it's never the bundle they are actually running in production.
Now, when I read some `sudo docker` I think the same things that I think when I read `sudo pip install`.
Plenty of software is distributed as "copy/paste `curl ... | sh`" or "npm install ..." or "pip install ...". This is absolutely not unique to containers.
> most of this software won't even run without root privileges
I don't buy this at all. The container runtime probably needs root privileges, but individual containers rarely need privileged access. Moreover, in many (all?) cases we can use security policies to prevent root containers by default.
> Running this stuff without any sort of checking and as root is bonkers, but it is what it is.
Again, true of any software, containerized or not. For what it's worth, I'm pretty sure people are more likely to inspect a docker-compose.yml than they are to decompile an ELF binary.
> We're kind of back in the Windows 95 era of packaging software as far as server backends go. Maybe it will change after some very serious worms and viruses his the Docker ecosystem.
We've always been in that era. The only difference is that today our systems are designed with more security in mind.
Ha, little do you know. It's common to bind-mount various system directories or UNIX sockets into the container. Also, does it matter when you're running a full OS inside the container anyways?
Hosting providers is a tiny slice of the pie, most Docker users are simple end-users looking to run random internet software. (E.g., Docker is the only way to install third-party software on LibreELEC, a simple media center OS for the living room TV.)
Sometimes I feel bad about our security posture and then I read stuff like this. Thanks.
it’s not that common, in production systems anyway.
> Also, does it matter when you're running a full OS inside the container anyways?
Containers famously don’t include an operating system. They use the host’s kernel.
> Hosting providers is a tiny slice of the pie, most Docker users are simple end-users looking to run random internet software. (E.g., Docker is the only way to install third-party software on LibreELEC, a simple media center OS for the living room TV.)
I don’t believe this is true. I would wager that the overwhelming majority of containers are running in the cloud or in data centers.
What protocol were you thinking of?
I really think a lot of criticism of containers is absurdly low quality (e.g., criticizing containers for issues that are universal to all software)--it feels like people are really grasping at straws. One gets the distinct impression that some people have spent years or even decades perfecting bespoke, rube-goldberg-esque application runtime environments and now containers are obsoleting their value proposition. Of course, I'm very hesitant to psychoanalyze and would never argue that any individual is so motivated, but this is the impression I get in aggregate.
Privilege minization is much harder when stuffing everything in a container. I'd wager that running Chrome normally is probably safer than running it inside Docker, for example, because not all sandboxing functionality works when running inside a container.
So it would depend on what software, and what type of container.
These things force the application to run in restricted environments where capabilities are reduced, system calls are filtered, apparmor profiles and selinux labels are applied, among other things.
These are the same sort of things that chrome, for instance, does to make it harder to do bad things from the browser. They are hardening techniques. Images are automatically checksummed on download, and for certain things even signatures are checked (more of this soon).
Just because they do not protect you from a kernel exploit (also actually not true, seccomp filters have prevented a number of kernel exploits from inside the container) does not mean they are not added security.
The important thing to know what your attack surface is and what is acceptable risk for the workloads you are running. And yes, it is important to understand that containers do not isolate you completely from kernel exploits.
It is also important to understand the control plane (runtime up to the orchestrator) is much more likely to introduce security issues than the container itself.
There are multiple independent projects involved in securing a standard orchestrated docker style container (some of the set of Linux kernel/Linux distro/runc/containerd/docker/k8s) and no obvious owner of overall security configuration and problems.
we've seen examples of this, e.g. k8s disabling Docker's seccomp filter, or more recently the difficulty in how to handle clone(3) and seccomp filters.
For me it's that comparison with dedicated security sandboxes, is that in other projects there's a single team handling the whole security picture , which is likely to make things easier to manage.
There are a ton of security mechanisms that are enabled by the ecosystem itself, even if it does introduce new complexities and does have certain weaknesses against full hardware virtualization. It also has significant and meaningful security strengths (namely in availability via lower resource usage) against exclusively using hardware virtualization.
This characterisation is too charitable, without qualifications. It depends a lot on the container runtime and configuration. Out of the box with Docker you don't get AppArmor or SELinux, and apps run as root where they normally wouldn't (because uid remapping, aka user namespaces is disabled by default, and people rarely bother to set up non-root users in containers).
As a bonus, applying security updates is left as an exercise to you, meaning it often doesn't happen.
Also, complexity is the enemy of secutity. Containers are yet another added layer that you have to juggle in your head when trying to make sense of the whole system you're building and operating. If used wisely and with good understanding, monitoring and processes, they can be a net positive despite this, but not necessarily so.
And docker at least has an update delivery story (same as microVMs as well) compared to traditional ops where there are patching cycles and anything that can't be yum/apt updated is basically ignored and updated on quarter/year time. Build a base image in your CI pipeline, update it every night, have your app build against it, smoketest deploys, and largely forget that patching exists.
They can’t. Not one developer I have worked with in the last 10 years has lifted a finger in the name of security.
This is why managing containers is a full time job by itself, a specialised discipline.
If you can’t afford an FTE to manage containers you can’t afford containers.
I don't recall why we don't turn on selinux when it is available.
Say what you want about kernel security, but most side channel attacks are statistically using syscalls which are highly unusual for a regular, hosted application to make. Using a combination of SecComp and LSM apps (SELinux/AppArmor) you can defend against most of these attacks.
You are right that containers alone are not sandboxes, but namespacing does help in terms of isolation. If you want to sacrifice performance for some further security you can use an application kernel, which further defends the host kernel. Additionally, you can try to sandbox with host and node level isolation for services, which refines the application syscall profile to be very consistent and predictable. Then if something unusual occurs on the host (like writing to disk) then you can take definitive action like shutting the host down. That's some of the principles Bottle rocket was written on, among others.
Microsoft ignores security reports related to container escape vulnerabilities because they don’t consider it to be a security boundary!
Every time there's a container escape or privilege escalation, it's treated as serious and patched. I look forward to the day when we can confidently say that containers are in fact a security boundary.
Edit: Please tell me down voter how to do this with Docker..
Edit: Lol @ comments. ...could mean... I don't think anybody talks about orchestrating containers without meaning it includes scheduling and so on. What Kubernetes, Nomad, Service Fabric and so on does. The tools we use. We can include Docker Swarm for courtesy if you want...
Edit: Nothing in the article mentions Docker Swarm you sillies..
If you want to understand things at a deeper level, you have to cut through the cruft. You can put linux distributions into tar files and run them without Docker (see systemd-nspawn). You can use cgroups to run ordinary Linux binaries. You can make a group of processes run on a fake private network. You can build OCI-compliant container images without ever running a single thing in a container. Docker conflates all of this and confuses people. Don't be confused, realize that "Docker" does a lot.