Self-hosting a high-availability Postgres cluster on Kubernetes
ryan-schachte.com
ryan-schachte.com
Is there any way to overcome hidden complexity biting your hand other than studying k8s extensively?
I would not advise trying this for a "side project at work".
Generally we all agree that we move to the cloud and it's "fully managed". If it goes down - that's the price we pay.
If you have to ask the question "what if the RDS goes down", then you are really in a different universe.
That last guarantee of uptime requires a ton of work, testing, and money, because you are all the way up and to the right on the curve of diminishing returns.
It does go down though, don’t neglect the possibility because it likely will happen. With very average workloads, I’ve seen RDS databases restart unexpectedly, read replicas being completely out of service, and even databases being completely frozen (can’t even connect as root).
I’d still go with managed, but it certainly doesn’t give full reliability :) you still have to consider “what if it goes down” - it will!
When you have 1500+ databases these cost add up. At that point, this kind of techniques are required to self host the databases. Price per DB with HPA na VPA is way lower than what you would pay for managed databases as well as you can hire a full time devops+dbadmin and still be cheaper.
Not when it's a core part of your business...?
RDS will fail, but it likely will come back up without much action on my part. K8s will fail, you will spend untold human hours figuring out the k8s failure modes, before figuring out the database failure modes (which likely are quite straight forward). It's just a cost that is not worth it.
In Kubernetes you "solve" a lot of application complexity by abstracting it into things like Helm Charts or Operators. It's fun and easy to use operators to deploy databases, monitoring, mesh, minio... This does not make complexity disappear, it just abstracts it by creating another layer with an interface (k8s API) you may be more familiar with.
Whole Kubernetes in practice is adding more and more layers on top of it. All those layers/tools are usually amazing but soon you need some real k8s skills AND some app skills to handle even unavoidable scenarios like ... upgrades.
That's why it's so confusing sometimes and has such diverse opinions on. There aren't many environments that run vanilla cluster...
As a relative novice in the space, I'm grateful to hear someone say this out loud. K8s seems perfect to me for quickly scaling transient stuff like pipeline workers, web servers, but I've always been pretty leery of giving up the trivial snapshotting and rollbacks and other creature comforts of old-school virtualization when it comes to deploying long running applications, databases, and so on. And I've always felt kind of kind of guilty for not being on board to just mindlessly k8s-all-the-things.
Why? it's an ancient Windows Client/Server app. Each "node" manages its own state and communicates with each other in this proprietary, janky-ass way. It takes 10 minutes to start up a node.
K8/Swarm isn't going to do squat for this team except maybe launch dev/test environments a little easier.
You can do that too with k8s with APIs which support more than just one backend.
The zalando postgres-operator also mentions as a feature:
> Live volume resize without pod restarts (AWS EBS, PVC)
What API doesn't support it? k8s has support for resizing PVs and has for a while now. AFAIK all 3 cloud providers (and more) support increasing the PV using their storage class.
> K8s is not really for stateful systems, yet
Is this written somewhere or is it just your opinion?
I suggest you to read CloudnativePG documentation as well as this article: https://www.cncf.io/blog/2023/09/29/recommended-architecture...
Also watch the video of my talk at last Kubecon in Chicago about handling very large databases.
I hope this helps.
I have been a very happy user of CNPG even with occasional issues (database backup to GCS tripped me few times, but it works - mostly a bit of UX that I never was sure wasn't some fail of mine).
Now I only really need to add some automation for handling "recover the database and switch over clients to it" that is more automated (I understand why CNPG doesn't do recovery to existing database, but it is a bit annoying)
Modifying IOPS and other volume attributes is something less frequently needed but we just released alpha support for that too, if you must need it.
We have also added support for reporting volume usage in CSI specs, which I know some operators use to automatically resize volumes when certain threshold is reached (I however do not recommend using ephemeral metrics for automating something like this). But point is - you can actually define CRDs that persist volume usage and have it used by an higher level operator.
Another thing is - k8s makes it relatively easy to take snapshots which can be automated too and that should give someone additional peace of mind if something goes haywire.
Obviously I am biased and I know there are some lingering issues that require manual intervention when using stateful workloads (such as when a node crashes), but k8s should be just as good for running stateful workloads IMO.
Another thing is - k8s volumes are nothing but bind mounts from host namespace into container's namespace and hence there should be no performance penalty of using them.
People use complexity as a boogieman to justify throwing together their own really wild chaotic & only so-so tested "simple" alternatives all the time.
To me, this feels like a modern wonder. We have layers of responsibility. Many people operate Kubernetes clusters already for all kinds of reasons. It provides a powerful broad base. Now on top of that, we can operate Postgres, with a very smart failover system that's super well tested & broadly used, that leverages this competent starting place.
What are the Fears Uncertainties and Doubts you have that make you scared about composition? What would help address specific concerns? To me, this division of responsibilities & use of consistent platform for a variety of needs feels like a huge win.
People love "simple" options but they're not. Run naked through the woods like savages option has appeal, but just getting started keeps adding up:
Sure, just add some bash scripts for some wal backups that go off-site. Easy! Install pgbouncer like one does, just a quick install, point it at the right systems. Setup some replication. Install and configure more kind of hairy software to make it HA. Configure some TLS cert yourself to not send naked traffic over wire. Add monitoring!
Then operationally, how quickly do you think you'll be able to fail over (with ogbouncer staying ok), failback, do an upgrade, add more replicas, lose a replica? Can you rotate cert reliably in a timely fashion? Can your replacement? How well did you document everything? Will you have all the monitoring you need when incidents start coming in, or did you just spitball a couple metrics into place?
The alternatives, in my view, are obviously bad. You can do them. Either cheaply with risk or industriously with effort & applied-talent. But having cohesive autonomic systems at our back that try to help, that can faultlessly do many common tasks with perfect accuracy (across unimaginable numbers of systems, with perfect consistency, in record time): that feels like a massively better place in the universe, one that I don't get why so many people kick scream & drag against. Rarely are their arguments well elaborated ("scary"), and their counter-suggestions feel like they massively underrated how multi-faceted & carefully connected production systems are, for good reason, and how hard it can be to remember to not forget to change X when you do Y.
Installing the Postgres on the computer, or on a VM.
Really, just the fact that some people keep asking this question is enough to question everything else they say. The alternatives are obvious.
I know, I did all of that.
Which is why I punt that effort to a k8s operator these days, because at the end of the day it does everything I did manually plus makes it easier for me to spin a new database that has WAL shipping and backups to different location.
Sure the dbaas might have superb autoscing and reliability. It may or may not have good OpenTofu/terraform providers available, or other custom niche deployment tools one would need to pick up & adapt.
Same questions as roughing it hacking together pg yourself, except now you have no power. Is observability going to be up to snuff? When performance problems come, how are you going to feel when you discover that analyst techniques a and b work fine but c and d can't be done because you are on managed service that doesn't actually give you full access to the db?
Costs can also add up, and are hard to predict. If you have a managed service, you'll be using a lot more network, which often has some cost (also some latency too). This isn't for everyone, but Kube makes it easy to start with anything (ex: RPi) and scale out. Works with hosted hardware with some local SSD, or I can get a 16-core 7950X and a bunch of gen 5 nvme drives and a fraction of a terabyte of ram for <$2000 and that'll scale me to godlike levels. Whereas I'll keep paying for something hosted.
I think it's so so ideal to have a reliably capable universal autonomic computing later underfoot. Most of the backend has had special specific answers to each question, each problem. We have been disjointed. Kubernetes provides a platform applicable to a vast range of workloads, where yes you need to tune and build different clusters for different workloads, but where the skills transfer much more readily. Everyone in Kubernetes knows how to ship resources definitions/manifests, and that base knowledge is all you need to learn to start consuming services - any services - on Kubernetes. Skill transfer is immediate. Having universal patterns (the API server), backed by autonomic operators making & keeping these resources real: it's so much better than the thousand different paths road of technology we've been on until now.
No. Technology is hard. Operating services is complex, and gets more complex as the scale gets bigger.
Any attempt at making stuff simpler is usually either moving complexity elsewhere or making it more expensive (and in some cases, both).
Upgrades is handled with an operator, and happens by waiting all queries to finish, draining all connections, and restarting the pod with the newer version. The application can connect to any pod without any difference.
I perform upgrades twice a year, never really worried about it, and never had any availability problems with the database, even when GCP decides to restart the nodes to update the underlying k8s version.
Adding Swarm or K8 just seems like redundant and unnecessary complexity.
That seems the crux of it.
The only "docker graph" stuff is in OCI container data itself.
So in both situations you have the same IO limits of block storage. The question is does k8s persistent volume api add enough of an IO bottleneck to cause issues, IME that isn't the case.
Now if you want direct attached NVMe drives for higher IO than a network attached block storage will give, then it might be easier with a VM vs k8s, but I can't speak to that much.
Features like k8s Operators are essentially professionally-written code that stands in for a human agent that would take actions such as managing DB nodes, renewing certs, and performing backups. If you use mature operators, you can save a lot of money for a bit of effort.
OperatorHub is currently the main resource I use, but GitHub stars aren't exposed in the search so I have been looking at the "Capability Level" chart and checking for Github popularity when I find one with the feature support I want.
I'm facing this exact same issue now when trying to find an operator for Redis. I am not sure if I am just missing out on the "right" option by limiting myself to Googling and Operator Hub and looking for the one with the most Github stars, so I am open to tips.
I don't know if all of these alternatives use statefulsets but I remember several doing so.
I've personally found cnpg to be pretty robust, and supports everything you will eventually need once you're locked into a solution (eg. Robust backups, CDC, replica clusters).
I'm yet to find anything of a similar standard for mysql.
EDIT: it should also be noted that CrunchyData is a proprietary solution and requires a license to use in production. This is not particularly obvious from their docs.
Statefulsets have their place but are surprisingly inconvenient for database workloads.
HA didn't handled this case at all whole cluster went in crash loop
there was also issue of huge pages caused crashing and not easy way to disable those without some dirty injecting of config files at runtime
there could be some my fault at misconfiguration on by side, but I wasn't able to figure anything better from docs
Hope it would be interesting for you.
[1]: https://ongres.com [2]: https://stackgres.io
after a lot of tweaking I made it so that the pods would be as big as the underlying hardware nodes. one pod one node. once that was working I realized that I was using the wrong tool for the job.
the kubernetes tooling added nothing but complexity. needless to say I let it run like that having had wasted about a week getting it to work
At work, we moved from hosting elastic search on bare VMs to kubernetes. By leveraging scheduler policies we are able to pack / over-provision ES node pods of different clusters onto the same Kubernetes nodes, allowing for far greater resource efficiency, while being able to handle node failure while maintaining availability across all clusters. Additionally, this simplified operations significantly as we can now leverage the operator to do cluster wide operations (e.g rolling restarts, node OS upgrades, ES version upgrades, etc...) fairly easily.
We did, however, go 6 years (and several hundred million users and trillions of documents indexed) without needing to use Kubernetes!
We will blog about this at some point this year.
Nomad is generally speaking a much more appropriate technology when you hit the point of needing such a system.
At the high-level you are describing your setup, it doesn't make sense. You'd spend way more resources managing any cluster than what you would gain from a normal-looking policy. I seem to be missing some important detail.
Things literally have just worked.
We are about to do a scale up operation of about 6 nodes each for 16 clusters (total of 96 ES data nodes being added.) This was editing a variable in a config file, and pushing that config to kubernetes. The operation will probably take a day to complete as we throttle the speed at which data transfers to new nodes to not adversely affect latency or ingest rate.
The amount of "human time" to kick off this operation and begin the scale up is measured in minutes, then it's just letting it do its thing.
Zalando is the company. ”Postgres Operator” is the software.
Happy user here, not much complaints about the operator come to mind.
Database works via this. S3 via minio. Redis. And maybe openFAAS?
That seems like a lot of the key building blocks already.
https://kiwiziti.com/~matt/wireguard/
Of course, there are tradeoffs you have to make (security, uptime, criticality of the system), but homelabs exist as perfect experimentation frameworks.
At $DAYJOB we run a global scale Ceph cluster (S3-like API/object storage), so even at a large scale it's not impossible to imagine.
Probably the most hard thing I found was wrapping my head around the way the storage works.
That being said, we're also layering a bunch of stuff on top - helm, nginx, GKE, terraform, as well as a mountain of other things, and then to top it off we have a bunch of shell scripts doing random things to help tie it all together.
Normally I can pick things up pretty quickly. I just built a parser with tree-sitter, despite knowing virtually nothing about language design. Didn't take very long.
But the modern devops stack is a learning curve like I've never seen before. It's taking me more energy to learn it than it did for me to learn programming itself. Then again, maybe I'm just getting old.
I resonate with feeling incredibly dumb whenever I pick up a new ticket from our backlog. It feels like gaining deep knowledge of these systems will be a nearly insurmountable challenge. It's been eight months, and while I know far, far more than I did on day one, I feel that every day is a day one of sorts.
Four times across three companies I have run across frameworks which set everything up out of band and then ask the tests to test it. Maybe these are shell scripts written by the k8siest person on the team. Maybe they're something like tilt. Whatever the case they're always a black box to the majority of the people who are writing application code.
They get you 90% there, but eventually somebody wants assurances that some environment variable has the desired effect, and suddenly you need to penetrate that black box and change it so that there are multiple kinds of "up" and the right tests run against each state.
K8s tooling is commonly installed via curl, so once you unravel the black box and integrate it with your tests you end up with a lot of fragile interfaces to things like kustomize, kubectl, kind... Fragile because maybe the other dev has a different version installed. Nix dev shells solve this, but you can't usually get the whole team on board with Nix so version mismatches come up often and are often difficult to debug. You end up in a state where whoever wrote the initial setup scripts is authoritative about the dependencies, and you have to ask them what they have installed if you want things to work (it was easy for them, they just used whatever was lying around at the time).
These aren't directly deficiencies of k8s, once you see the light (which takes a long time) it's pretty easy to work with, but like so many other technologies, the devil is in the peripheral tooling and the culture. K8s doesn't (yet?) have a very nice boundary with other language ecosystems, it reminds me of Java in that way. The die hard k8s people often want to solve problems by bringing them more fully into the k8s way of seeing the world and I just don't think that is consistent enough with reality to be our everything.
I've cultivated a begrudging respect for it, but I still don't like it. If I break free and start my own company, I'll publish an operator so that my stuff can be installed into k8s, but I don't intend to make it primary in any way.
If you don't need much orchestration (which is true for a lot of postgres users), the complexity from kubernetes is compounded on top of the complexity from postgres without generating much value.
If you learn it from the top, immediately jumping through bunch of deployments, Helm (ugh) and other stuff?
You'll get fast to deployment but you won't know how to deal with things failing and there will be a lot of stuff that will remain "magic". Eady way to end up cargo culting despite best intentions.
Can be enough of you're just making simple apps to run on it but not running the clusteror the application in production. But you will have hard time understanding why things work and everything will be complex upon complex.
Go from the bottom up, learn basics of kube API patterns, kubelet, how scheduling works (don't have to be in depth), how kubelet works, how networking works (CNI, why kube-proxy is a bandaid for applications that handle networking badly, how services work), how storage works (how kubelet mounts things to containers, etc). Then how higher level controllers (aka operators) work - from Pods, through ReplicaSet to Deployment, StatefulSet, DaemonSet.
This way you'll learn the basic building blocks, which are quite simple despite the long list I just gave, because the architectural and API patterns repeat and build over each other. The core is simple which lets you build complex stuff on top while still understanding it.
You should test your setup, because it's very often not done right, and it's easy to overlook a problem.
Obviously it's worth doing simulated disaster recovery to ensure you would recover if there is hardware failure. The larger the scale and throughput with parallel writes against the same keys, etc then the more complicated the setup will be. I hope to write more on this topic, but setting up a persistent volume with a NAS is a great way to ensure high durability.
Text files?
Creating != Supporting
Anything can be supported by creating your own driver.
OP uses Longhorn which is a whole other thing that I've only read about.
For at home you can use other storage classes like ceph, NFS, etc.
Cloud provider provide managed versions of Postgres that are highly available. Even if you self host, Kubernetes isn't the answer.