Why I'm building a Home Lab with K8s on Raspberry Pis
iamsafts.com
iamsafts.com
It's cheaper, in my opinion, to get a refurbished mini desktop from Lenovo or whatever. They're "real" PCs, and I've found the raspberry pi hardware limitations to be onerous. It's just really nice to have SATA, NVMe, and PCI Express lanes/slots. For the price of a couple of PIs, you can get a powerful CPU with loads of RAM.
Also i think you can use ryzenadj to set a very low tdp and disable the fan?
My HP EliteDesk 800 G4 has an annoying fan even at the lowest speed. It rattles. The newer G6 model I have at work is quiet.
A few weeks ago I picked up an ex-enterprise server with 256GB ECC RAM, 24/48 Xeon cores, and a handful of disks - for AUD$1000 (~ 600 USD)
Yes, it's a noisy box, and yes 2.5" SAS disks are expensive & small - so you want storage separately - but the same applies to a RPi or NUC based lab.
It draws about 100W while idle (and in a home lab environment it takes some effort to make it sweat)
Proxmox installed and runs like a dream.
I'm sticking with Nomad over k8s (or variants) for container orchestrator across several Debian VM's in the one box -- so I'm effectively relying on ECC, dual-PSUs, and hardware RAID5 for my ersatz HA.
I won't be doing any 3d printing or anything like that though, so what the author here is doing looks fun in it's own right!
How is your Longhorn performing? I tried setting it up on my nodes but with Gigabit networking and possibly the Pis pretty average CPUs I would get pretty awful performance both in distributed volumes (with replicas etc) but also on strict-local ones (for reasons I haven't yet figured out).
I am now considering using something like https://github.com/rancher/local-path-provisioner, since I mainly intend to use Longhorn for DBs that handle fault tolerance / backups etc on their own.
Volumes become unmountable, pools stay yellow forever, desynchronization happens... Perhaps (most probably even) I just don't know how to configure them properly, but maybe that is telling of the difficulty to operate/maintain these solutions.
What I took away from this is to try my most to never deploy anything that requires shared persistent volumes. If I need something stateful, it needs to speak to a database or S3 backend, or something that handles redundancy at some other level than filesystem. If I really really need to have a local volume, I'll use local-path-provisoner like you said, which means pinning the pod to a single node, but really that is a concession I am willing to make to not deal with ceph/rook.
Great writeup, it's like I'm reading about my own journey managing kubernetes clusters!
Best of luck
Things I would do differently are using NixOS or bootable containers (CentOS) (side note, bootable NixOS container would be a killer app) and writing my own helm charts instead of fully customizing my manifests and doing the deployments from ansible, and would recommend against Raspberry Pis for the compute as the 3's and 4's don't support limits, e.g. cpu or ram limits, and I wasn't able to set up firecracker containers correctly on the Pis.
I'm also exploring hyperconvergence infrastructure (HCI) as that seems more like my ultimate goal for homelab stuff.
With enterprise gear it becomes outdated and then you replace it at your leisure, whereas with consumer equipment it dies so you need to replace it. The disvantage of that is the noise and power consumption so it's a tradeoff you need to consider.
One interesting thing you can do is run PBS as a giant VM and back up or migrate the whole thing, just like you can run a small NAS as a VM as well.
Wasn't aware of NixOS, looks pretty interesting but I'm not sure about how easy / reliable it'd be to run it on a Pi 5 (https://wiki.nixos.org/wiki/NixOS_on_ARM/Raspberry_Pi_5). I'll be keeping an eye on it though!
As far as Helm vs Ansible, I'm using Ansible to deploy the basics (bootstrap control plane & worker nodes, networks plugin) and then everything is deployed with IaC (Pulumi) which installs Helm releases.
How do you like Pulumi? It seems similar to AWS CDK...
Most services I've been using so far offer official Helm charts. But I get your point, it can be cumbersome and if there isn't an official one, then they can be pretty undocumented / hard to work around.
I haven't used CDK, but the concept is definitely similar. I think Pulumi most likely has wider support, since it's based on Terraform and even if you don't have a provider available on Pulumi you can "port it" (although never tried it, not sure if it works well). I like how it stores the state for you and secrets as well, saves quite a bit of trouble.
There’s a lot of irony saying this about bare metal vis-à-vis k8s. Maintaining individual servers isn’t difficult, doing it at scale with high availability and the requirement of upfront investment (hardware, colo, staffing, etc.) is. Doing k8s at scale isn’t a walk in the park or a cheap date either, though.
But, as mentioned, I only see it as a developer. My end product with k8s resembles something that's closer to my development tools than what I'd have maintaining a series of servers and using other tools to deploy apps onto.
Cargo culting the processes required for extreme use cases outside of that context is only good for learning to get paid money to do it in those contexts in the future. In and of itself it is very silly. For a human person a process running on an OS on running on a computer is far, far less work, maintenance, and complexity with better performance and longer lifetime.
Eventually this system probably will become quasi stable, sustaining. And for some folks, replicable & practicable. The author will probably soldier through a good amount of the pain then it'll tick along. Doom isn't certain. (But it is probable!)
Such a process can be deployed/backed up/restored in minutes by anyone with basic sysadmin skills and can scale to serve nontrivial user bases while costing much less than SaaS.
Classic sysadmin is seriously underrated.
I tried doing a homelab with six Nvidia Jetson Nanos using K8S, maintained it for about a year, and I have no desire to ever do that again. I ended up just buying a single rack mount server and using that for two years, and now I bought a mini gaming computer which I use as a single server. Maintaining the k8s cluster was becoming a second job that actually costs me money, that I enjoyed less and less every day, and it made me dread actually using any aspect of my server, meaning that when something broke it would take me a long time to actually muster up the strength to fix anything. My home server runs a Transmission server, Jellyfin, Apache Kafka, Apache Spark, Cassandra, and RabbitMQ, with 32 gigs of RAM, and it works fine.
Distributed systems are cool and they're fun to play with, but the combinatorial explosion of maintenance shouldn't be underestimated. If you're making something that needs to serve 10,000+ users, then it's probably worth it, but homelabs generally aren't that. Generally a homelab situation has like a Plex/Emby/Jellyfin server, a torrent server, a reverse proxy, maybe some kind of message queuing solution, and it generally only has like four concurrent users.
Obviously if your goal is to learn K8s, then doing it with a bunch of Raspberry Pis isn't a bad idea at all, and try to have fun doing it. However, I would warn anyone thinking that they're going to make their NAS a k8s cluster, you're likely going to regret it. I recommend buying a slightly beefier computer and just installing NixOS or something.
If you get frustrated and don't want to do it anymore, you might also look at Docker Swarm. For homelab stuff, I found it considerably easier to work with. I'm not sure how well it would work with hundreds of services, but for a homelab I found it considerably less frustrating to deal with than K8s. It can still be useful to play with distributed systems and it's a bit more intro-friendly in my opinion.
It's something that I think most engineers should try at least once, if nothing else to learn firsthand the annoyances of distributed computing. It's easy to read about these things in textbooks and blog posts and university courses (and you should!), but at least for me these things didn't really sink in until had to fix broken stuff, or deal with weird heartbeat issues, or see things that should go N times faster actually go slower because I didn't take into account latency etc.
I opted for bundlewrap over Ansible/TF/etc, and keepalived over MetalLB.
I also didn’t add any kind of real storage or networking: they’re just using WiFi.
After I got it set up and working, and was forced to read up on the high-level k8s concepts, I never really did anything much with it again, but the learning was valuable. :-)
kinda annoying that on this blog post, if i press the up arrow cursor button or page up button it takes me to the top of the page , instead of scrolling up
the page down and cursor down button work as expected
--
kinda found the bug, when you first open the blog, the cursor stays on the top site navigation menu, as long as the cursor is up there, page up or cursor up will take you to the top of the page
click anywhere on the page that will change the cursor position, and now page up and cursor up work as expected
I have some hardware for fun and learning, but it is well isolated. Using K8s for home infrastructure is bad idea.
With single server, complexity is much much lower. I installed my server years ago, once a month I check logs and do updates (10 minutes of work). Some time next year I will upgrade to Ubuntu 24.04 and replace SSD (just in case), about 1 hour of work.
K8s is really horrible when it comes to time efficiency!