Talos: Secure, immutable, and minimal Linux OS for running Kubernetes
talos.dev
talos.dev
https://github.com/siderolabs/talos/issues/8367
Unless I’ve missed something, this isn’t a big deal in an AWS-style cloud where extra storage volumes (EBS, etc) have essentially no incremental cost, and maybe it’s okay on bare metal if the bare metal is explicitly designed with a completely separate boot disk (this includes Raspberry Pi using SD for boot and some other device for actual storage), but it seemed like a mostly showstopping issue for an average server that was specced with the intent to boot off a partition.
I suppose one could fudge it with NVMe namespaces if the hardware cooperates. (I’ve never personally tried setting up a nontrivial namespace setup.)
Has anyone set up Talos in a useful way on a server with a single disk or a single RAID array?
This is one of the more annoying problems with any "immutable" distro (or at least the ones I've tried). They demand a very specific partition layout that basically forces you to devote a whole disk to them, which as you said in a cloud environment (or any other environment where VMs are used in place of physical machines) doesn't matter, but ends up mattering a lot when using physical machines.
Are you referring to things like a home directory?
On an edge-style machine, you generally have a very limited number of M.2 slots, and dedicating one to a boot device uses up potentially 100% of your potential storage capacity, not to mention that an extra drive would be a non negligible fraction of total cost.
On modern servers, there may not be a usable SATA controller, and NVMe usually costs 4 PCIe lanes. Those aren’t free. You also end up paying the absurd OEM premium for an extra disk if you go with a big name OEM. Sadly, there is no industry standard cheap, reliable boot device standard, at least as far as I’ve ever seen. Maybe someone should push USB3 for this use case — the price is certainly right, and performance is likely just fine.
Also, an external dangly thing is asking for trouble (getting dislodged). An internal device solves this.
In any case, the Talos people seem to recognize this as a problem and are working on it.
There are also SATA-DOMs or USB-DOMs. You could use one of these modules with your machine.
To accomplish this, I have restricted myself to SQLite as storage and use Litestream for replication. On start, Litestream reconstructs the last known state before the application starts. [Source](https://github.com/LukasKnuth/homeserver/blob/912cbc0111e44d...)
It works very well for my workloads (user interaction driven web apps) but there are theoretical situations in which data loss can occurr.
Of course, if you really lean in to a complete lack of local persistent state and you configure your network or some other critical service like this, good luck recovering from an upgrade and complete loss of ephemeral state.
For my personal usecase this is fine. I also have monitoring setup to look for just this case. It's a tradeoff between resilience and simplicity that works for some use cases - mine included.
Facebook supposedly got locked out of their own datacenter due to a network outage preventing the access control system from accessing whatever service it needed to allow anyone to open the door.
The upside of my solution is that there is no scheduling requirement on which node the PVC was initially created. There is also a certain guarantee that I have a working, recent backup of the application data. Starting from scratch every time is also basically a backup recovery operation. It gives me confidence that there is a recent backup which is restorable.
But if I read this correctly there shouldn't be an issue adding one blank disk to the Talos VM, the issue is only more granular disk and partitions management.
Something that Talos does differently is everything is an API. Machine configuration, upgrades, debugging…it’s all APIs. This helps with maintaining systems way beyond the usual cloud-init and systemd wrappers in other “minimal” distros.
The second big change is Talos Linux is only designed for Kubernetes. It’s not a generic Linux kernel+container runtime. The init system was designed to run the kubelet and publish an API that feels like a Kubernetes native component.
This drastically reduces the Linux knowledge required to run, scale, and maintain a complex system like Kubernetes.
I’ve been doing a set of live streams called Talos Linux install fest walking new users through setting up their first cluster on Talos. Each install is in a new environment so please check it out.
Before that, we had a Kubespray based setup. It's a bunch of Ansible script and it allows to make any custom setup, like absolutely anything as you in control of the machines. But the other side of this is that it's extremely easy to break everything. Which we did a couple of times. And so any upgrade is a risk of loosing the whole cluster, so we decided it must be run in VM with full backup before each upgrade. Another problem that it takes about an hour to apply a change, because Ansible has to apply all the scripts each time.
Then we migrated to Talos, and it's a day and night. The initial setup took like an hour, including reading the docs and a tutorial. Easy to setup, easy to maintain, easy to upgrade (and it takes minutes). Note that we run the nodes as VMs in Proxmox, so the disk and network setup are outside of Talos scope, as well as backups, and it's actually simplifies everything. So it "just works" and we can focus on your app not the cluster setup.
Talos improves security further by mounting the root filesystem as read-only and removing any host-level such as a shell and SSH.
After host-level, probably 'access'.
And there is a repo of them here - https://github.com/siderolabs/extensions
They’re even shareable images. Eg here’s the image with gvisor included https://factory.talos.dev/?arch=amd64&cmdline-set=true&exten...
https://www.youtube.com/live/HsY8D9aO84Y?si=VL5LPG_M9GwfM7d_
Talos doesn’t support older models (too slow) or the 5 yet (waiting for uboot support)
When you get sick of patching let us know
Also you shouldn't really use pod OS to debug it. Kubernetes supports debug containers: you launch a separate container (presumably with convenient debug environment) and mounts selected container rootfs inside, so you can inspect it as needed. It also helps, when the target container does not work and you can't just exec into it.
There's a recommendation to remove everything from the container that's not necessary for running a given program, that reduces attack surface.
Correct: Your dev environment should also not let you do stuff on the host machines. In an k8s environment, you run everything in pods. Don't compromise on security and operational concerns just because it's a dev environment.
> If you can't login to it then it is not good for development.
You develop inside pods, and you are more than welcome to install any shell and other programs you want inside containers. (Or for working at the k8s level it doesn't matter; you `kubectl apply` or run helm against the k8s API, it doesn't matter what's happening on the host.)
I guess we'll never know...
But in truth it's for running on hosts.
So, being dogmatic about "the host should not have any tools installed" is good and all, but how do you debug this scenario without tools on the host?
We eventually figured it out. By logging into the host OS and using the shell tools there.
What was the cause/solution? Images too big?
The cause was indeed images being too big. Images — not only the raw images, but also their extracted contents on the filesystem — count towards ephemeral storage too. In their case they can't even control the size of the images because those are supplied by a vendor.
The solution was to increase the node's disk space.
Less dogma, more the lived experience that letting people log into hosts ends badly. Though I grant there's a cost/benefit both ways and perhaps there could be edge cases.
> but how do you debug this scenario without tools on the host?
Cordon the node, evict any one pod to free up just enough room, and then schedule your debug pod with a toleration so it ignores the error condition? I confess I've never had to do this but it seems workable.
In this time, I remember having to SSH into a host node exactly once. This was me, the platform engineer - not an application developer. Even then, having is a strong word. I could have just as well done with a privileged container with host access.
Application developers have nothing to do on the host. As in, they gain nothing from it, and could potentially make everything worse for themselves and the other applications and teams on the platform.
> Production ready: supports some of the largest Kubernetes clusters in the world
> It only takes 3 minutes to launch a Talos cluster on your laptop inside Docker.
> delivers current stable Kubernetes
Whilst you're not wrong, and the website could be clearer, there are plenty of clues.