Kubernetes StatefulSets are Broken
plural.sh
plural.sh
K8s has really been an experience of encountering 2-4 year old bugs (closed by the stale bot of course) and missing features. We accumulate increasingly more workarounds and additional pieces of complexity to deal with this. Everything is always changing, but somehow still staying the same.
* You can create a volume, but pay only for in-use blocks.
* You create a volume bigger than what you'll ever need => no resizing required (At least 16 TiB, but possibly even something crazy like an exabyte)
* There is some kind of "trim" operation that marks a block as free
* Ideally it's possible to choose the maximum block index (size of block-volume) independently from the maximum number of in-use blocks (for cost control)
Every business is incentivized this way. Competition is what pushes prices down, of which there is plenty in the cloud industry (though maybe still not as much as we would like).
Edit: Now that I actually read TFA I see it mentions most CSI drivers already provide this functionality. They provide a workaround similar to what I’m sure most people use, and I agree this functionality seems like it could be handled by StatefulSet.
I guarantee you internally cloud providers over provision their storage.
since then, I'm totally on board with pre-provisioned storage. at the very least, some way to put an upper bound on volume size (maybe efs has that? not sure.)
Unsurprisingly, AWS EFS does not support NFS quotas.
There is always a point where you can tell in advance that it doesn't make sense to keep going. It is never sensible to literally go on forever with anything -- it can only break stuff in annoying ways when it runs into real world limitations (like financial ones in that case.)
I sound aggressive about this because it's such a common mistake. It always goes something like,
"Why does this list have to be unbounded?"
"Well, we don't want to give the user an error because it's full."
"Okay, but does it really need to support 35,246,953 instances?"
"Sure, why not?"
"How long would the main interaction with the system take if you stress it to that level?"
"Oh, I don't know, at that level it might well take 20 minutes."
"And the clients usually timeout after?..."
"5 seconds."
"Would the user rather wait for 20 minutes and then get a response that might be outdated by that time, or get an error right away?"
"They may well prefer the error at that point."
"So let's go backwards from that. Will it ever make sense to support more than 250,000 instances?"
"That corresponds to the five second timeout and then some. I guess that's fine in practise..."
It's not that hard!
Often limits need to be addressed at the product level. At some point at my $work we started pushing hard enough back at product to say that we're placing a hard limit on every entity. What the limit is can be negotiated, re-evaluated, and changed -- but changes to it need to be intentional and done with consideration to operational impact.
My personal pet peeve is when i'm dealing with a client library that doesn't expose some sort of timeout.
I'd also add unbounded queues to the list above as a subtle place where the lack of a limit can really cause production issues. Everything may look like it's working fine until you realize that you've got a 4gb process with a giant buffer that'll take forever to drain.
here it is: https://news.ycombinator.com/item?id=32175328
No snapshots. No support for ACLs. Poor performance. High cost.
SAN vendors have been selling devices that deliver thin provisioned network block devices you can snapshot, clone, resize dynamically and replicate in a cluster for about 20 years now. Third parties can do all of this in cloud environments as well. See Netapp Cloud Volumes. You could roll your own using Ceph, or even a ghetto version with a subset of these capabilities with thin provisioned LVM+iSCSI, DRBD etc.
You should be able to point and click a thin provisioned, dynamically expandable, snapshot-able, clone-able, replicate-able pay as you go network block device with whatever degree of performance, availability and capacity you wish to pay for. The fact that you can't is a function of the business model and mentality of IaaS cloud operators; they don't want your old fashioned 'stateful' workloads. Neither does kubernetes. So you're swimming against the current.
I haven't used it, so your mileage is a total unknown, but our "big data" folks like it.
AWS won't change to your desired model, they mint money on underutilized resources (and hidden I/O charges/network)
Practically, sizing thin-provisioned filesystems too large wastes a lot of space on filesystem metadata.
All operating systems and clouds that I've worked with support online volume and filesystem resizing which is as close to the semantics of thin-provisioning as I've ever needed. I wasn't actually aware that StatefulSets don't resize volumes as smoothly as desired but I haven't had a problem manually deleting pods to get the filesystem on a PVC to grow (the way it works for Deployments).
However, I've always been suspicious of operators, so maybe my bias is leaking here. Kubernetes is already a level of abstraction and adding another seems risky and unnecessary.
the whole point of code as infrastructure is that your git repo is an audit log of whats changed in your infra. Once you leave it up to operators or persistent volume claim templates (or whatever else makes side effectful changes with being told) you're throwing all that niceness of control out the window.
Right now what i have is "oh the cluster is in a different state (there is drift) than the code, lets be a detective to see why this changed".
So you start debugging the helm chart, and kustomize, and the operator.
Not fun.
You can't make a baby in 1 month with 9.5 women.
It was painful. We quickly realized that rolling upgrades often got blocked. The reason being peristent volume claims sometimes get stuck.
And this is just one of few problems.
Since we moved databases and stateful services outside Kubernetes, everything got faster and more reliable.
I personally haven't experienced the issues the OP is describing. It might be that StatefulSets and PVCs are more stable now then when OP tried them. Of course using managed/hosted K8s makes a big difference.
There might have been a way to support our use case better. But this was managed service and we couldn't easily tweak every settings. But also, we didn't want to.
Our goal of using Kubernetes is to make difficult stuff easy. Not to make easy stuff difficult. And in our skillset and capacity moving stateful services outside K8s was easier, cheaper and more reliable.
K8s is great way to easily deploy and scale services. As we leverage Kafka, a lot of applications fallback on Topics and KTables for storage.
Kubernetes with jobs, cron jobs, operators, secrets etc is effectively an Operating System for the cloud.
Definitely worth the investment to learn and practice.
The fix was to add scheduler hints that moved the pods to the correct zone as I should have done in the beginning. (On first deployment, the disk is created on the correct node and thus zone, so it all seemed to work).
Just curious - how does GKE control plane upgrade break DNS?
The cluster's DNS server refuses connections briefly during control plane upgrades or updates. I haven't tried their node local DNS yet (enabling it will cause downtime, natch), maybe it'd help.
Either way, this is a good example (for me) of how Kubernetes needs some work to be friendlier to operators
That said there are warts that should be removed, but that's not surprising.
You might have drained requests off of a pod, but it still has a "state", it still has things cached in memory, it might still have connections open to some outside entity (Database, whatever), and the developer, not K8s, is responsible for catching those signals, cleaning up things in memory, gracefully terminating connections, handing off in-flight workloads to another pod, etc. Even in the upper echelons of the tech I see a very small minority of developers actually aware of all the things that can make stateless workloads stateful. Which is OK for something where the stakes are low, but if you're a DBA or a Systems oriented person you'll see people make these (very wrong) assumptions all the time and recoil in horror.
Theoretically is not actually, and a ton of the assumptions in k8s break down when you look at them super closely, but they break down when things have gone wrong. If your cache is properly designed it doesn't matter if it's a little stale; you're not using it for things where correct and up to date is the most important thing. Connections, too; it shouldn't matter if your connection pool needs to spin up another connection, because if any of the conditions exist that the connection fails (or takes too long and stalls requests for any appreciable time), you already should have been paged and be on the way to resolving the incident.
If the PVC size changes Kubernetes automatically does an online resize of the filesystem: https://kubernetes.io/blog/2022/05/05/volume-expansion-ga/
This has been possible for 2-3 years if you had the flag enabled.
StatefulSets are valuable for applications that require one or more of the following.
Stable, unique network identifiers.
Stable, persistent storage.
Ordered, graceful deployment and scaling.
Ordered, automated rolling updates.
[1] https://kubernetes.io/docs/concepts/workloads/controllers/st...
When do you not need automated rolling updates? Or stable network identifiers? I'm sure there are cases, but it seems like supporting them should be the default -- are statefulsets somehow expensive?
Stateful apps can be a bit more exceptional because they often have client libraries that require you to hardcode network addresses via connection strings and other tooling that makes that glue automation more tricky and worth the added guarantee around consistent naming.
Most of the time with stateless apps.
In regards to automated rolling updates, the key word there is "ordered." It will always select the same ordinal index first(reverse order starting with highest ordinal.)[1] This is different than the regular deployment strategy of "rolling update" which is also one by one but will select any pod to whack first.
Regular pods don't have stable network identifiers as they are inherently unstable. In regular pods stable network identifiers are provided by either the cluterIP or a load balancer service.
[1] https://kubernetes.io/docs/tutorials/stateful-application/ba...
And I would guess the doc is a bit of after the fact justification for a design choice already made.
I believe that's only partially true, since even if one has a PVC in something like a Deployment (the antipattern I see *a lot* in folks new to k8s trying to run mysql or whatever), then "k scale --replicas=10 deploy/my-deploy" will not allocate new PVCs for the new pods
That's why I think StatefulSets are valuable as their own resource type: they represent _future_ storage needs
It’s really not that big of a deal and probably a ten minute process, tops. That said, OP makes a good point that mechanisms to handle this kind of stuff exist in CSI, so why not leverage / take advantage of that versus pissing in the wind (especially when you run into CRD that wants to undo the above work while you’re in the process)?
Now that it has, it’s really a matter of someone having the time and patience to drive it though. I don’t know whether someone has sorted out all the design implications (like what happens if someone adds new resource requests, or what happens if a new resource request is rejected, and how that gets reported to the user), but it’s a small change that needs a fair amount of design thought, which tends to be hard to drive through.
Even from the early days, there were constructs to allow stateful applications to be built on top of it.
They have made it easier to run stateful applications, but to say it wasn’t designed to support stateful applications is incorrect: https://kubernetes.io/blog/2016/08/stateful-applications-usi...
“Stateful applications such as MySQL, Kafka, Cassandra, and Couchbase for a while, the introduction of Pet Sets has significantly improved this support.”
I have been using Kubernetes for a few years now and whenever I see people putting manual steps they should instead be using an operator. What the author is describing as a problem has a solution (they even added it to their own operator).
Kubernetes has a steep learning curve but if you want to do nontrivial thing like have state than you need to learn nontrivial things like operators.
The default controllers that come with Kubernetes do not cover everything under the sun and they do require extensions via the operator pattern for some tasks.
Just my two cents.
So when a k8s node dies with a sts on it, unlike other pods, the sts pod is NOT recreated (not wanting to interfere with whatever mechanism the db cluster has to deal with it). I've found most people working with k8s don't know about this behaviour.
I have a feeling that OpenShift handles stateful workloads more gracefully due to their enterprise-oriented customer base, but I have nothing to back that.
Also, for important "big data" stores with replication, I don't want to trust all-in-one third party operators, which is what all K8S operators seem to aspire to be. Your data is important, and you should evaluate all your data operations use cases yourself.