Ansible-Defined Homelab
0xc45.com
0xc45.com
Please help me understand your choices. You went for reproducible VMs on your fourth, custom machine. The three NUCs are compute-only (stateless?) nodes for a k8s cluster.
What kept your from making your custom-machine a node as well? That gets you Infrastructure-as-Code too, since your Nextcloud (what I am setting up right now, too) and whatnot are defined in code, using containers.
By the way, since my set-up is a single NUC, I ended up going for docker-compose. That has just plainly worked. Single-node k8s with outside access (only need HTTP/HTTPS for Nextcloud and the like, with reverse-proxying through subdomains for more apps, like CodiMD) didn't prove to be friendly to me (not my field of expertise). I tried had tried multiple CNIs, traefik and contour as IngressControllers, and MetalLB as a LoadBalancer, but could not get it to work.
And even if it ended up working, it would have to be a stateful k8s cluster, and I imagine that adds a whole layer of problems, too. Do you then go for NFS to the Synology NAS? Seems like the best idea, but that will be much slower than the nodes' local SSDs. Do the SSDs then sit idly, without work to do? Seems like a waste. If the deployments/PVCs/PVs are tainted such that they are bound to specific nodes, you don't have much of a cluster, more like a spread-out docker-compose (which still beats docker-compose, I guess, if it gets running).
These are the questions I couldn't answer, or whose answer drove me to docker-compose.
Since k8s has so much more steam than docker-compose, I imagine in 1-2 years, k8s will be the go-to for single-node homelabs as well. What do you think?
Not gonna happen because it takes 3 master nodes to have a quorum and 1 separate worker node. The requirements are also pretty high with 2GB of memory per virtual machine. Don't think people have 4 extra machines or 10 GB of free memory to run VMs.
Can you explain that more? I had a single-node cluster running: the master needs to be untainted so work gets scheduled on it, and then it runs. There is no need for a quorum because there is nothing to decide on.
Kubernetes depends on etcd which requires an odd numbers of instances. There might be ways to run a one or two nodes cluster for testing (see minikube) but that's not how it's going to run in a company.
Most companies won't be running their own etcd clusters anyway, except for on-prem clusters. Cloud users will generally use a managed Kubernetes cluster. If you do need to learn how to run etcd, you can learn the basics in about a day. In our org, the etcd portion of the training is less than half an hour long thanks to good documentation.
For a homelab, consider running something like k3s which can use a SQL database instead of etcd.
Some Kubernetes-based distributions, like OpenShift, require 3 masters at minimum, but that's not the norm across K8s installations.
For example, I picked up about 96 GB of DDR3 ECC ram for around 75 euros. A quick check on ebay, and the same amount of DDR4 is selling for at least _twice_ as much. I imagine it's pretty economical to buy this older hardware and just assemble a single beefy server, instead of buying multiple physical machines.
The added benefit is that this older hardware doesn't end up in a landfill, and even though older CPUs generally consume more power than their current gen equivalent, I reckon a single machine would consume about the same, or less, than multiple NUCs (I have no source to back that up though, it's just my assumption).
I agree the NUC is a great platform, but if you could spend less cash and get more bang for your buck, and perhaps have the added benefit of having a platform with ECC memory (not sure if the NUC supports ECC, I'm assuming it doesn't), then I think the latter is what most people would go for (or well, at least what I would go for :p).
We're also talking about home _servers_, so it doesn't seem that odd to me to use actual server hardware. The homelab[0] subreddit has a bunch of folks running actual server hardware for example.
I had plans to build a noise isolated data closet in the basement (tied into the furnace air return) but I never ended up with the right sort of basement.
You don't, strictly speaking, need separate worker nodes. You just need to untaint the masters.
Or you can just run the masters on raspberry pis - under $200 for a properly redundant cluster. (Masters and workers do not have to share a CPU architecture)
Management processes shouldn't take 2GB. What are they doing in there?
Very important thing to know about memory: Kubernetes nodes have no swap. Kubernetes will refuse to install if the system has a swap, gotta remove it.
This means nodes better have a safe margin of memory, because there is no swap to use when under memory pressure (things will crash). Hence the minimum requirement of 2 GB.
I tried to run complete clusters with 5+ machines in VmWare, with as little memory as possible because I don't have that much ram on my desktop, and all I got was virtual machines crashing under memory pressure.
Previously I used hand-crafted, Ansible, Salt, Puppet and Docker setups (that last one died two weeks ago prompting me to start the K3s setup). But they all ended up becomming snowflakes, making them to much hassle to maintain. What I like best about the K3s setup is that I can just flash K3OS[1] to a SD card with near zero configuration and just apply the Kubernetes resources (through Terraform in this case, but Helm is also a great timesaver).
I still have to figure out a nice way to do persistent storage. But for now since it's one node (the IoT cluster does not have state) the local-path Persistent Volumes work well enough. Might have a look into Rook.
I will admit it's not trivial to get started with Kubernetes, but since I already needed to study it for work this provides me with a nice training ground. Alternatively Hashicorp offers nice solutions in this space as well with Consul and Nomad. But it needs a little bit more assembly, whereas K3s comes with batteries included (build in Traefik reverse proxy, load balancer and DNS).
https://github.com/kubernetes-incubator/external-storage/tre...
MetalLB is dead simple to set up, and you can hook it up with a DNS server with TSIG to get automatic name entries for your ingress hosts.
About the custom "VM-only" machine: currently I'm a little lacking in my "k8s ops" skills. So, until I am more comfortable managing the k8s environment, I have decided to continue using VMs for my "production" workloads. Until then, the K8s cluster will be treated like more of an experimental zone and I can worry less about breaking things.
As for the K8s storage options -- I don't really know! Haven't quite figured that one out. Like I mentioned in the blog post, I haven't really placed any serious workload on the K8s cluster yet. However, I have heard good things about OpenEBS (https://openebs.io/) and am considering that as a potential option in order to make use of the local SSDs on each NUC.
For single-node K8s, there are lots of good options. I haven't used it personally, but I know that MicroK8s (https://microk8s.io/) works well even in a single-node environment. It might be worth checking out.
I'm running my kubernetes nodes as VMs and I didn't want to have to backup the virtual machine disks. I just wanted to back up my freenas server, so I went with trying to use NFS.
I wanted to use the NFS client provisioner, as I thought I'd want to be able to add and remove volumes automatically, however, this makes redeploying onto a new cluster from scratch hard. Eventually, I just tore that out and just mounted each NFS volume individually because it turns out that I actually only need a handful of volumes.
However I see that it has multiple backends (including CEPH), might be worth looking into as it seemed very fault tolerant and fast. :)
It was a lot easier to setup and had more of an "it just works" out of the box experience compared to openebs/rook/etc.
About building your own NAS to replace the synology, take it from someone who has done this, don’t bother.
While rather easy to do, Synology provides so much more “out of the box”. My own NAS ran well on Debian 10 for a few years, until I started getting random disconnects on drives and various other “stability” issues. Synology costs more up front, but after initial configuration it’s more or less just a box in the corner you forget about. Mine just sits there, automatically installing software updates whenever they’re available.
As for my own current setup, I run my hypervisor (Proxmox) on a Dell PowerEdge T30. It runs an internal docker host (Debian 10) and an external FreeBSD host on a DMZ vlan for anything that is accessible from the internet. Everything on that box runs in its own jail.
All storage except OS storage is mounted from the Synology via kerberized NFSv4.
Firewall is handled by a Netgate SG-3100 running Pfsense, but I’m in the process of migrating to a UniFi Dream Machine instead. Pfsense has been good to me, but the SG-3100 is expensive (for what it delivers), and its beginning to show its age. I have a 300/300 mbit connection, and with Suricata enabled I get frequent reboots because it uses too much CPU, causing the watchdog to think its stuck. The UDM is half price and twice the hardware, and for my usage (router/firewall/vpn/dns blocking) the UDM does it all (Pi-hole through a 3rd party solution).
Does your synology utilize BTRFS? If so, what is your experience with it? I'm currently using a FreeNAS box that runs on ZFS that has been very reliable for me over the years. I can't really see giving up the power/reliability of ZFS to move to a synology setup, but I like to keep tabs on BTRFS progress.
As for the firewall, I had a Ubiquiti Security Gateway for a while and it worked great. Not sure how it compares to the UniFi Dream Machine, however. The only reason I made a custom firewall is because I wanted to try and learn a bit about firewalls and routing in general. And, since my own device has worked well for me so far, I haven't bothered to replace it.
I can't easily create new VMs, create snapshotted/linked clones, and I can't easily pull out CPU/RAM/disk utilisation information from ESXi for the hypervisor and all VMs to store in a time series database for monitoring/alerting purposes.
There's also a tonne of additional features in Proxmox that requires additional licensing for ESXi (and yes, I am aware of the VMUG EE licensing making a lot of that a lot more affordable)
Proxmox ultimately being Debian based gives a lot more flexibility in that regard. It's also a potential weakness as a result, depending on perspective.
Maybe since people are suddenly spending more time at home, there are certain resources that they'd normally have access to at their work that they don't have access to anymore and need/want to create themselves, or are suddenly more reliant on their own infrastructure, or may just spending more time at home leads to interacting with their home set-up more often which leads to more tinkering with it. I know I've put much more effort into working on my home dev set-up this year than I did last year.
I got 4 bay Synology NAS with Intel CPU & 8GB RAM which nicely runs Docker, some reddit users boost it even up to 16 and 32GB RAM. Plus OpenWRT router. Entire "rig" eats like 40Watts.
The only reason I can see having a VM as an abstraction layer is if your hypervisor runs in a cluster and it's able to move volumes across nodes, but even then, it feels like having a single VM running on the nodes that take up the full physical resources makes more sense.
Prioritising time for this kind of thing seems to go in cycles for me. E.g. i remember spending a few weeks back around 2005 setting up a home lab “just so”. I still have the notebook i wrote up at the time. Loads of knowledge in there i’ve since forgotten about setting up lvm on aix hosts with smitty and pinouts for converting cat-5 & db9’s into null modem cables for managing cisco ios on old switches. I can barely remember config t and enable mode these days...
There’s a real satisfaction to be had from this kind of thing. Serial consoles on everything, jumpstart configs dialled in just right to be able to rebuild any node on the home network hands off.
Then it was around 2013 the next time i got an urge (the ansible repo above) then most recently it’s been k8s experiments. Interesting it’s about 7 year cycles.
My ideal solution would be to create configuration for a VM and its applications and run a command to have the VM created, OS installed, applications installed, etc. Taking it a step further, being able to regenerate my reverse-proxy configuration, certificates, and DNS configuration would make lighting up new services incredibly easy.
Right now I'm running my VMs on FreeBSD and managing them with vm-bhyve. There are possibilities for automation here, and I've done a little bit of it (at work) using Kickstart to install CentOS VMs and shell scripts to do the rest. Unfortunately, this is all very purpose-built - if I were to do it again, I'd probably scrap it in favour of Ansible or something similar.
Obviously I've got a long way to go to get to my ideal.
For me, that line is drawn at hand installation of a particular OS and then conversion into a template. From then on, I can reference that template in Terraform or Ansible automation and clone it before customising it to the requirements I have that day.
The only requirement from using these tools against VMware is a vSphere instance. I’m sure if you look hard enough you can find that quite cost effectively on eBay.
And just remember that your infra is never finished. It’s a journey! And the journey is often more fun than the destination! Happy labbing friend.
For example, I can go to Digital Ocean or Oracle Cloud or Azure or AWS and get a VM provisioned in a minute or so. I know some of the tools involved, like cloud-init, but I'd really like to know how the rest of it works and to see it in action.
Honestly though, the number of times I need to build a new VM is very small. The time savings that come with each individual VM build will never recoup the time spent building out the pieces required to make it happen. But knowing how that all works will be worth it.
I also run OpenBSD HA firewall clusters with trunking (aggr), VLANs at home and work:
- https://github.com/liv-io/ansible-roles-bsd/tree/master/host... # interface configuration
- https://github.com/liv-io/ansible-roles-bsd/tree/master/open... # firewall configuration
https://github.com/liv-io/ansible-playbooks-example/blob/mas...
Admittedly, this is unlikely to be an issue at home lab levels. Therefore one must suspect the “because I can” methodology is at play here.
Not that I dislike LXD, but a good configured server won't require many interventions. In my case, I was about to upgrade some guest SO, and I had to relearn lots of things and remember what I did, even why, after about two years without touching it.
Even if it’s just for me I still try to write reasonable commit messages and all that.
Personally I think the amount of pain wasn't enough, so I'm using Canonical Maas, and I can just spin up extra servers, though going beyond 4 nodes seems kind of wasteful in a home setting. Though my rig seeps 100Watt, which for hobby purposes is okayish imo. I've been running this since Ubuntu 16.04 and upgrading to 18.04 has been a relative breeze compared to a situation where I would've built this using adhoc scripts(i.e. non-k8s setup.)
I did make a wireguard tunnel to a colocation datacenter for offsite backups using async mirroring of rook/ceph. B/c with so many moving parts and automation, wiping a cluster accidentally is quite possible.