In this time, I remember having to SSH into a host node exactly once. This was me, the platform engineer - not an application developer. Even then, having is a strong word. I could have just as well done with a privileged container with host access.
Application developers have nothing to do on the host. As in, they gain nothing from it, and could potentially make everything worse for themselves and the other applications and teams on the platform.
> Production ready: supports some of the largest Kubernetes clusters in the world
> It only takes 3 minutes to launch a Talos cluster on your laptop inside Docker.
> delivers current stable Kubernetes
Whilst you're not wrong, and the website could be clearer, there are plenty of clues.
Also you shouldn't really use pod OS to debug it. Kubernetes supports debug containers: you launch a separate container (presumably with convenient debug environment) and mounts selected container rootfs inside, so you can inspect it as needed. It also helps, when the target container does not work and you can't just exec into it.
There's a recommendation to remove everything from the container that's not necessary for running a given program, that reduces attack surface.
Correct: Your dev environment should also not let you do stuff on the host machines. In an k8s environment, you run everything in pods. Don't compromise on security and operational concerns just because it's a dev environment.
> If you can't login to it then it is not good for development.
You develop inside pods, and you are more than welcome to install any shell and other programs you want inside containers. (Or for working at the k8s level it doesn't matter; you `kubectl apply` or run helm against the k8s API, it doesn't matter what's happening on the host.)
I guess we'll never know...
But in truth it's for running on hosts.
So, being dogmatic about "the host should not have any tools installed" is good and all, but how do you debug this scenario without tools on the host?
We eventually figured it out. By logging into the host OS and using the shell tools there.
What was the cause/solution? Images too big?
The cause was indeed images being too big. Images — not only the raw images, but also their extracted contents on the filesystem — count towards ephemeral storage too. In their case they can't even control the size of the images because those are supplied by a vendor.
The solution was to increase the node's disk space.
Less dogma, more the lived experience that letting people log into hosts ends badly. Though I grant there's a cost/benefit both ways and perhaps there could be edge cases.
> but how do you debug this scenario without tools on the host?
Cordon the node, evict any one pod to free up just enough room, and then schedule your debug pod with a toleration so it ignores the error condition? I confess I've never had to do this but it seems workable.