Kubernetes Cordon: How It Works and When to Use It
cast.ai
cast.ai
Pretty nuts to think I’ve saved six figures in VM costs just doing that once or so a month.
Or, think of it as defragging.
You basically have enough work for 3 nodes but it’s spread over 5 nodes with some spare capacity on each.
Kubernetes is not killing the extra workloads to consolidate it all on the 3 nodes it would have fit on; kubernetes only does bin-packing on new workloads.
The cluster-autoscaler component does remove nodes with low utilization and Pods fitting onto other nodes, effectively defragging the cluster. There's just plenty of reasons for it to not drain a node (as per its FAQ[1]).
But it's highly configurable, so I wonder if the grandparent post just needed some configuration changes to suit their scenario better.
[1] https://github.com/kubernetes/autoscaler/blob/master/cluster...
I can’t be sure which one the GP is referring to, but GKE definitely operates the way they described.
My point is that while the original post is reporting problems with autoscaling in Kubernetes without enough details to go by, cluster-autoscaler is fairly good at its job of downsizing clusters if you take the time to a) optimize its configuration and b) make sure your workloads are configured in a way that allows downsizing clusters.
[1] https://cloud.google.com/kubernetes-engine/docs/concepts/clu...
Silly default IMO but apparently "intentional".
Draining a node will allow it to shut down.
Of course, if you want to keep the node up and running, but evict the pods, then absolutely that is when `cordon` _should_ be used.
If you want to drain 3 nodes as fast as possible it's best to start by cordoning all 3 of them and after that run drain on each respectively. This will cause minimal unnecessary interruption of pods. What we want to avoid is that new pods start on nodes that we soon want to remove/drain.