Understanding AWS End of Service Life Is a Key FinOps Responsibility
fairwinds.com
fairwinds.com
For one, depending on your situation/CRD's/automation, doing these upgrades in-place can be next to impossible. Updating an EKS minor version can only be done one version at a time - e.g., if you want to go from 1.24 -> 1.28, you need to do 1.25, then 1.26, then 1.27, then 1.28. So teams without a lot of resources are probably in a tough spot depending on how far they are behind here. Often, it's far more efficient to build an entirely new cluster from scratch and then cut over - which seems ridiculous.
Why are upgrading EKS versions such a pain? Well, if you're using any cluster add-ons, for one, all those need to be upgraded to the correct versions, and the compatibility matrix there can be rough. Stuff often breaks at this stage. Care needs to be taken around PV's, the CNI, and god help you if you have some helm charts or CRD's that rely on some deprecated EKS API - even if the upstream repository has a fix for it, you will often find this yak-shaving nightmare of fixing all the stuff that breaks on upgrading that, and then whatever downstream services THAT service breaks - etc.
What is the solution? I don't know. I'm not a kubernetes architect, but I work with it a lot. I understand there are security patches and improvements constantly, but the release cycle, at least from an infrastructure/operations perspective, IME places considerable strain on teams, to the point where I have literally seen a role in a company whose primary responsibility was upgrading EKS cluster versions.
I have a sneaking suspicion this is to try to encourage people to migrate to more expensive managed container orchestration services.
Realistically those workloads being run would have been better suited in an horizontal-scaling EC2 deployment but that was a future goal that never came to fruition.
For reference, I have done upgrades from 1.12 -> 1.28 and most of the time if get into a messy project and I can get away with it, I will just rebuild a cluster from scratch.
Yesterday I spent 3 hours trying to fix something and find it's an indent error somewhere.
If K8s would be backwards compatible, upgrading would be a lot easier, and if they would support LTS releases, like other projects, manual upgrades would be needed only every X years.
For example, the reason that you can use PostgreSQL with the same major version for 5 years on RDS is due to PostgreSQL actively supporting it, and minor versions are non-breaking and can be seamlessly applied (restart or failover to standby replica is still needed during upgrade).
Having new realeases so often for such a core infrastructure component is kinda insane unless it was explicitly architected to allow seeamless upgrades.
The hairy bit is the rando junk that gets shoved into clutsers, without any sane packaging scheme to roll it up or back. I even recently had to learn the deep guts of the sh.helm.v1.foo secret because we accidentally left an old release in a cluster which no longer supported its apiVersion. No problem, says I, $(helm uninstall && helm install --version new-thing) but har-de-har-har helm uses that Secret to fully rehydrate the whole manifest of the release before deleting it so when helm tries (effectively) kubectl delete thing/v1beta1/oldthing and pukes, well, no uninstall for you, even if those objects are already gone
The question is: what service? ECS is the competing Amazon built service and it’s entirely free for management, you just pay for compute. We don’t use k8s because ECS is free and we don’t plan on leaving AWS.
Sure, you’re more locked in with ECS, but if you aren’t doing funky stuff with the APIs you can probably off ramp to k8s pretty easily. I know I could move us in a week or less, we’d have far bigger problems with the other AWS services we use.
sure, I guess there's a whole industry that offers services to help manage AWS, but at that point the whole thing could be just spun off to the lowest bidder. (although I understand that people can make up a neverending list of reasons (excoughses) to keep paying the AWS tax.)
... okay, I'm probably too salty. if it works it works, if it's profitable and business is happy, it's hard to argue with the stack.
And for Kubernetes, honestly, charging 6x for extended support is probably a bargain, considering the pace of change and difficulty of hiring engineers for unsexy maintenance work.
We just rolled off of their extended version and it was about 19 minutes to upgrade the control plane, no downtime, and then varying between 10 minutes and over an hour to upgrade the vpc-cni add-on. It seemed just completely random, and without any cancel button. We also had to manually patch kube-proxy container version, which OT1H, they did document, but OTOH, well, I didn't put those DaemonSets on the Nodes so why do I suddenly have to manage its version? Weird
Touching CNI is always a potential downtime inducing event, but for the most part it was manageable
Haven't had that situation myself on AWS yet, but ran into it a few times on Azure
I can't remember to have paid extra on Azure though, but maybe we did. Certainly not 6x the price though.
PS: not sure why it got flagged the first time, but I think because I used a different title. Sorry.
Great example of misuse of that simple word 'that'.
Should be 'which'.
"The difference between which and that depends on whether the clause is restrictive or nonrestrictive.
In a restrictive clause, use that.
In a nonrestrictive clause, use which.
Remember, which is as disposable as a sandwich wrapper. If you can remove the clause without destroying the meaning of the sentence, the clause is nonessential (another word for nonrestrictive), and you can use which.
[1] https://www.grammarly.com/blog/which-vs-that/#:~:text=Which%....
https://www.youtube.com/watch?v=y8OnoxKotPQ
> Which we get from EKS -- our entropy chaos service.
Charging an arm and a leg for a deprecated service is another.