Kata Containers: The speed of containers, the security of VMs
github.com
github.com
After installation, it is constantly running. As can be seen by:
ps aux | grep docker
And it occupies IPs: ip addr | grep
I was on a train in Germany recently and could not use the Wifi because of that. Turned out the docker daemon occupied the IP range the Train Wifi uses. While I was not using docker at all.I guess it clutters even more stuff. Any suggestions what else to look for?
So I am looking for a cleaner container solution. One that feels more like a Linux tool that keeps the system intact and only runs when it runs.
If Kata is such a tool, I would look into it more closely.
Podman is what Docker should have been, for me. Security first, no daemon, more Linux-like behavior (You can manage them with SystemD unit files if you wish) and it supports the same, usual container images you build/use with Docker.
The main part it was lacking is the compose equivalent, but that too is coming along.
See my other comment in a recent thread [0]
Does it support higher-level declarations like Deployments and StatefulSets? I'm trying to understand how/if we could use this without having to write new manifests. A (very) quick search didn't clarify it for me.
All in all its a great program and IMO even better than docker but It would be great if people try to make it sound like its a 1:1 comparison to Docker because it has its trade offs.
Selinux is neat but is IMO conceptually wrong in a container world. UIDs and are barely better. At most, these mechanisms should confirm that the container as a whole has a given permission, and that’s it.
(Seriously, the major clouds have deprecated per-object permissions on their object stores. IMO they are right to have done so.)
Podman's use of SELinux prevents and mitigates many of these issues what you will have if somone targets your app.
the reality though is that you can install both docker and podman, and just start/stop docker as necessary. its easy enough to experiment with podman on a system with docker installed.
imo it's similar to folks that think learning a new shell is a huge undertaking. the reality is you just install the new thing and drop in and out of it as you get comfortable. if it sticks, cool, if not, also cool.
This would be the equivalent of me buying a great car and disliking the fact that when you pull the carpets up there's a hard to reach clip that needs to be undone, so I decide to sell it and get a different car.
I see this happen with all great tools. It irons out all the important kinks, and people still find some obscure reason to fault the tool enough to switch.
(To be clear, I am aware Docker has issues beyond what GP is dealing with)
(Apologies for the strong tone of the comment, it's not intended, but could not find a better way to word it)
There is nothing shitty about the network on the train for also using this same IP address range.
https://en.wikipedia.org/wiki/Private_network
10.0.0.0/8
172.16.0.0/12
192.168.0.0/16
You might be used to seeing addresses from 192.168.0.0/24 and 192.168.1.0/24 in home networks, and addresses from 10.x.y.0/24 in corporate internal networks.
But all of 172.16.0.0/12 has exactly the same kind of purpose as do 10.0.0.0/8 and 192.168.0.0/16.
The people that set up the network on the train did nothing wrong for using a subnet of 172.16.0.0/12.
Docker Compose creates a new network for every project, and eventually overlaps with something important. They are fairly large ranges by default, so you end up taking up a lot of address space fast if you're not careful. This is especially wasteful because some of the Docker networks only contain two hosts, but are (from memory) a /24 or maybe even a /20.
Easier to handle it manually.
That’s not a docker/container problem.
Btw, you‘re able to configure the Ip-space to use by passing an argument to the daemon or change it in daemon.json.
Also, yes, it is a docker problem in particular. Linux has a solution for virtualizing networking for a subset of processes only: network namespaces. Docker doesn't use them by default, but can be taught to do so with the rootless kit. All rootless container engines use them by default.
as long as ipv6 link local are not disabled the fe80 addresses work nicely :)
Using IPv6 would have reduced the probability a lot. But the excuse for the last 20 years has been, why bother with learning something new as long as "it works for me"... (I don't claim I would do differently.)
That aside, while I do agree that small private IPv4 space availability is a real concern, I'd also argue that Docker choosing to make the default network size a /16 compounds this problem significantly. I've never had a workflow where a Docker network needed more than a /24, and most could get away with a /26 or /27 without it being considered an aggressive limitation of IP space. Assigning the default Docker network size to something much more reasonable for a development context would do wonders for limiting collisions like this in the first place.
Configure daemon address pools in case they conflict/intersect with real network addresses.
Now its even muddier with Canonical ... taking back (?) LXD.
Throw in the alternative vision (LX* containers are more like persistent VMs than ephemeral containers) and a lack of `<container-engine> pull app` and all that entails re: DevX and DeployX, it always felt like a mountain to get going.
Is quite doable with lxc. An application container can be built and shared. It doesn't appear to be typically done, though.
Do you know any resources to learn it a bit more in depth ?
A few years ago, I made a proposal to have some automatic grafting mechanism: https://github.com/NixOS/nixpkgs/pull/10851
This would automagically work by simply maintaining 2 trees of Nixpkgs, one with the cherry-picked security updates, and one which matches the latest set of cached packages. This way one can fully benefit from the cached packages while having the ability to replaces with the latest security patches they want to import without building the world.
Unfortunately, rewritting Nixpkgs to fit the requirements needed to have the automagic mechanism is a huge project, especially given the activity of Nixpkgs. Maintaining a fork of Nixpkgs which stays up-to-date while changing its inner working cannot be held by a single person.
My hopes would be to push this to the Nixpkgs Architecture Team, while preventing them from doing mistakes by inserting extra complexity while making this work more challenging.
As for security, it's worth noting here that there are Nix-native tools for generating MicroVMs as well, if that's what folks are after with Kata and Firecracker.
Nixpkgs includes the same kind of reuse and integration and patching that you see in other kinds of software distributions, like Linux distros or Conda.
For reference, this can be worked around by deleting your Docker networks, logging in and recreating the networks, which should pose no problems on a dev machine.
I am not aware of any guidance how to use privat IPv4 addresses. In practice 192.168.1.0/24 seems to be the most commonly used one, so you might want to avoid that.
You need to understand what containers are first. Containers are not one thing. They are an amalgam of different OS primitives designed to give you the maximum flexibility, control, and isolation for an application environment.
When you say Docker is "cluttering" your system with processes, you mean the daemon that is used to start and manage containers on your system. There are alternative container systems that don't use a daemon, and can run rootless, but they have some tradeoffs. They are also not nearly as portable or easy to use as Docker, as a whole.
Yes, it "occupies" IP space, by default. You can disable or reconfigure the networking aspect of container solutions, to either use a different subnet, or just use the host's native networking. But then you won't get network isolation for your containerized app, and you will probably complain that you can only run one process on a given port at a time, and without a firewall, people on the train will be attacking your containerized apps.
> So I am looking for a cleaner container solution.
There isn't such a thing as clean software. People like to generalize like this, but what it usually means when they say "clean" is "I want it to be magic, as simple as possible, do everything I could ever want, and to not have to think about it". Which is wanting to have your cake and eat it too. Either it does everything for you and it's complex, or you have to get your hands a little dirty and it's simple.
> One that feels more like a Linux tool that keeps the system intact
Point in case: you want it to maintain the system for you.. Docker does that. The end result is what you call "clutter".
> and only runs when it runs.
You want a rootless daemonless container frontend, like Podman. Good luck getting it to work... Don't @ me when you find out it's a lot of extra effort that doesn't give you anything better than Docker did.
Kata containers is for service providers. Nobody really needs that level of isolation on their laptops.
I was surprised when I learned this but Docker by default bypasses UFW and potentially other firewalls relying on iptables.
https://blog.viktorpetersson.com/2014/11/03/the-dangers-of-u...
Yeah, that's the service you installed. Stop the service if you don't want it running.
You can also have containers use host-mode networking (share's the host's NICs instead of using bridged networking via virtual NICs) if local IP address pollution is a concern.
Kata is just a container runtime. Depending on how they implement their network and storage drivers, functionality should be mostly the same.
Has it gotten better? Any resources you recommend? Cursory Google searches on this have so much outdated info it can be hard to quickly wade back in.
Last time I tried it the standalone Docker/containerd integration wasn't working well, the project seemed to be more targeting deployment as part of a k8s cluster.
Is there a good reason why they don't seem widely adopted?
TBH their docs aren't that great. There should probably be a 'curl | sh' solution to install it at the top of the readme followed by a '<run this command and you're in an ubuntu shell in kata!>' command right after.
Another issue is the lack of nested virtualization in EC2 instances that aren't the very expensive i3 metals. That turns this from a "it's a drop in replacement" to "I'm spending thousands of dollars on this".
https://www.redhat.com/en/blog/red-hat-openshift-sandboxed-c...
Please don't pipe the Internet directly into your command line.
It gives a quite lot more trust than running arbitrary content as shell script, without any third party verification.
Which is still just code from the internet.
> , without any third party verification.
Certificate provider does not verify packages nor anything what is coming from there. Server even might be just proxy.
Of course, if you are a target of nation state attack, which fakes public keys from all sources by MITMn DNSs and servers, you might end up with the wrong package.
But that threat model is totally different.
Ummm, so what do they do again? Sounds like marketing speak. "Like VMs but containers" doesnt tell me anything
It's explained in more detail here: https://serverfault.com/questions/773581/virt-manager-partio...
There are other orchestration systems that can use LXC - LXD, libvirt, Proxmox, and may be others. Also, LXC doesn't have traditional virtualization - that's a feature of LXD using KVM. (Do you mean system containers, as opposed to regular app containers?)