1. Get a new server
2. Connect ethernet and IPMI ports
3. Add node definition to XCAT
4. Set boot to network, arm node for installation, power-on
5. Profit.
6. Get a coffee.
It's 15 minutes for [1-150] servers. Slightly longer for 150+. Because, network.Any further setup is done via Salt, if necessary. With a single state file.
Fun fact: I provision K8S nodes that way too.
apiVersion: apps/Betav1
kind: Deployment
to: apiVersion: apps/v1
kind: Deployment
Because in K8s 1.16, they stopped the backwards compatibility for Beta on deployments. But that was quite some time they left it in for backwards compatibility, and it was a simple sed script to change it. We have never changed our Dockerfiles except to change the FROM line when we update a base image. They often introduce new features, but try to keep things as stable as possible for existing setups.In our case, we store our YAML in git, and can deploy a new cluster with all our microservices in about 10 min (most of that is delays in the google global load-balancer setting up a TLS certificate for us, the machines are up and running after just a few minutes).
There can be a bit more scripting needed to do updates of the nodes to newer builds with no downtime, if your pods aren't totally stateless, but compared to what it used to be at an old job, with C apps running on Linux, behind load-balancers that had to be removed, upgraded, and added, a rolling update is like black magic.
This is exactly the bit I care about: what do I do with "stateful" systems (my db storage, mail server data,...)?
If I just mount those from external volumes, I lose a lot of idempotency and I now have to worry about whether I am trying to attach that volume to PG 9.3.1 or 9.3.2 or 12.0 (some are fine, others might cause data corruption).
I know idempotent deploys are all the rage, but all those deployed apps are there to serve some data which is as stateful as it can be.
I'd like a system where those external volumes are automatic snapshots on LVM (or ZFS/btrfs) when attached, but that introduces a whole another level of what-now if you need to go back to an older, now slightly stale data set.
How does k8s solve this problem for me?
These files can idempotently be re-applied to a cluster to restore, or upgrade, the running applications.
When you lose your cluster and you restore to a new cluster, you will of course initially deploy the same version of kubernetes as you were using previously.
Then once you choose to upgrade to a new version of kubernetes, it can happen that some of the APIs you were using have now been deprecated/removed. This is a very slow process, so you had plenty of time to upgrade ahead of time before being forced to.
But let's say you ignore all that, and have chosen to upgrade to a new kubernetes major version and are now forced to upgrade all your yamls. This happens rarely, and recently only because a few BETA APIs have become STABLE and people are now expected to use the STABLE versions. So you go ahead and make those few changes, re-run your deploy and you're done.
Kubernetes is still early tech and thus has more moving parts than a more well established stack. Keeping up to date is crucial and no small feat.