PostgreSQL 14 on Kubernetes
blog.crunchydata.com
blog.crunchydata.com
If we need to increase storage this can be achieved in minutes, if we need to take one node down we basically can do it anytime during the day..
Every now and then we discuss moving DBs to AWS RDS (or equivalent) however for our transaction volume RDS costs are prohibitive.
How can I know that this isn't one of those cases? I know some folks in fintech and while obviously they're not _you_, they're definitely not the sort to be wishy washy with data, if anything they're quite conservative when it comes to how the databases are managed.
Any incidental complexity is incredibly discouraged.
a) run your stateful services (such as Postgres) on the same orchestration system
b) choose a second orchestration system (often a more "traditional" infrastructure provisioning system like Chef/Puppet/Ansible/SaltStack) that will only manage your stateful services, and deal with the care & feeding & mental overhead of having two things that do fairly similar jobs and need to interact with each other (because of course your k8s-hosted stateless services need to be able to look up Postgres in DNS)
(option C would be to use a managed service like RDS...but options A and B assume you've already made a decision, for whatever reason, to run the stateful services yourself)
The main force that's pushing us to move our storage into K8s isn't simplifying our node configuration, but to keep parity between our prod and more dynamic staging environments. For the latter k8s storage operators are great because we don't generally care about data durability or performance.
Not quite true if you use a managed K8S offering - and it seems a self-managed cluster is something to be avoided until absolutely necessary. Even with a managed cluster like GKE there's plenty of operational overhead to keep everything up-to-date and working properly.
I'll be downvoted to hell for this.
What does Kubernetes offer in this scenario?
I'm still hesitant to run things like DB's on K8's but have absolutely no problems running cloud native apps on K8's which can run ephemerally, be bin packed and only needing compute as all state is external to where the app is being executed.
A kubernetes node can advertise those local block devices as persistent volumes. Your StatefulSet can request that class of local storage and the kubernetes scheduler will then find a volume for the pods on a node that the pod can fit on.
You then get all the features of kubernetes operators plus the performance of dedicated nodes and local storage. When a local storage node disappears, your pod gets a new volume on a new node and can replicate from its peers, or start from a backup.
Kubernetes then offers metrics, certificate management, secret storage and management, load balancing, healthchecks, debugging tools, rich firewalling, RBAC, and a huge ecosystem of plugins. It gives you a unified API for all your workloads. Requesting a database then is as simple as a single API call to kubernetes with a spec for your database (ex: creating a MysqlCluster object).
The company I work for runs MySQL, ZooKeeper, memcached, hbase, hadoop, kafka, elasticsearch, and more in kubernetes with great success. Scale here in one prod region is ~10,000 pods on a clusters of ~600 nodes. There are about 900 separate MySQL clusters consisting of 3 replicas each, 150 separate zk clusters of 5 replicas each, etc.
Actually no. Currently you need to manually delete the persistent volume claim. Followed by deleting the pod. For the statefulsets ti to create a new PVC.
There even was a race condition until like 1.21 that made this operation fail 20% of the time so you'd have to delete the pod twice or the system would deadlock.
It'd be nice if the PVC would claim a new PVC if the orchestration layer knows 100% the PV is gone but it currently doesn't do this yet.
It's not an enormous problem once you're aware of it. But it does mean things don't auto heal at all. A node going down with local storage on it needs manual intervention to get the system running and healthy again.
There might be some operators that automate this away but i haven't seen it yet in the wild.
Kubernetes is neat, but destroying stuff is so easy (and the context mechanism is just begging for human error) that I am a bit leery about hosting large amounts of stateful stuff in it.
PodDisruptionBudgets allow you to specify the minimum number of replicas. The drain command calls evict instead of delete and evict honors these restrictions. This is helpful for node draining, but not necessarily helpful if someone accidentally deletes a StatefulSet. However, nothing in kubernetes deletes PVC objects automatically, so if you delete the pod, the volume remains bound to the PVC until you delete it.
preStop hooks are used in all the database pods to ensure that if the current pod is the primary, leadership is correctly changed to a replica before proceeding with the delete. We combine this with extremely long termination grace periods and alerting if termination is taking too long so that an ops team can look into the issue.
preStop hooks are great, but they also are somewhat of a time bomb if not careful, so some of our database operators make use of finalizers instead. With finalizers, nothing happens when a pod enters the terminating state until an operator marks the pod's finalizer as completed. This flow is really nice because you end up with behavior where the kubernetes cluster admins or built in controllers effectively are really gently requesting that a pod terminate, and it's up to the operator of that pod to decide exactly how and when that occurs.
For non-local volumes, like EBS PVCs, once the PVC is deleted an operator will snapshot the volume during a finalizer before allowing the volume to be fully deleted so that we can recover data if this ever accidentally occurs. This was the first protection we ever implemented when initially adopting kubernetes back when we had no idea what we were doing so that we could always recover data in the event of admin mistakes. We can't do this for local volumes, however, any team making use of local volumes truly needs to stay on top of replication when running in the cloud, whether on kubernetes or not or else you absolutely will lose data. AWS local nvme volumes have limited write cycles just like the ones in on-prem servers. So ideally all databses are configured to not accept writes that cannot be replicated. And even with EBS, at scale we end up with something like 20 unrecoverable volume failures per year
Mostly I'm concerned about preventing administrator mistakes; if you have enough privileges, it's distressingly easy to just accidentally delete all kinds of resources in the default configuration, especially since kubectl is context-sensitive. I like my automation software to have checks in it to prevent me from doing stupid things without explicitly disabling a number of safeties first.
I feel like Kubernetes is a powertool that requires a competent and trained admin team or managed access so that you simply can't make awful mistakes as a "regular user", but a lot of the hype around it seems to be focused on how "easy" it is.
To my eyes, it's indeed easy to just throw manifests from the internet at Kubernetes, but that is often akin to disabling SELinux because some blog says so; it may get things working, but it's not the competent choice most of the time.
I'm also a big proponent of storing everything in git repositories, so I do like how Kubernetes enables declarative configuration, though it seems much of the tooling in the ecosystem works against that by just creating resources in a cluster with no version control...
The additional protection we have is via admission controllers which hook all API calls to prevent mistakes, of course, you have to foresee those mistakes and come up with a validation that prevents it. But even something as simple as "only 5 pods are allowed to be terminating at once" can help.
I absolutely agree about people claiming kubernetes is easy. It is easy to throw some helm chart at the kube API and suddenly have an app running with all it's databses configured. It is hard to keep them all running.
The pattern of operators seeks to solve this. Instead of a helm chart creating a MySQL statefulset, it creates a MySQLCluster that gets managed by a well-written operator following the best practices and correctly implementing all the hooks. The ecosystem is starting to converge on this finally and the beauty is that the sum of the industry's operational knowledge can be coded into these operators and we'll finally arrive at nearly fully self managing databases with backups and all.
The industry isn't there quite yet though, and so at least for stateful workloads you need significant operational experience. For background: I am the Tech Lead of our team that runs the kubernetes clusters and we provide all of the primitives our users need and control access. Each database type then has dedicated teams operating them and codifying operations into operators. There is a lot of work being done. But prior to kubernetes, all of these teams separately were coding against the EC2 API with no standardization, no unified view, ad-hoc failure detection, ad-hoc auditing, etc. Kubernetes is a very substantial improvement over that, but 90% of companies never reach a scale where this is necessary. But once extremely solid open source operators exist we may truly hit the dream of just applying some self managing operator manifest to an EKS or GKE cluster and getting an actually production grade database setup.
Version control is an interesting topic because it's hard to express transitive dependencies in that code. For example, the MySQL operator creating a StatefulSet and the Statefulset creating pods. It doesn't make sense to commit those lower level resources to git, they're not fully independent. However, for top level things, our build system produces artifacts from git and the deploy system creates them in the cluster. With this setup, only admins could directly apply them, which really helps with keeping things in sync with git.
My ideal is that what is in the repository allows me to rebuild the system from scratch (assuming suitable hardware exists) to the point where it can be successfully used as a restore target for your backups following documented restore procedures.
- StatefulSets mean that K8s scheduling is aware that the workload is stateful and the auto-attach of drives works fantastically in my experience.
- NodeTaints/Tolerances and workload anti-affinities between workloads mean that you can very easily say “this workload runs on nodes with these features/labels” and “don’t run workload replicas on the same node” and “don’t run other workloads on the same node as this workload. Node groups on things like AWS EKS also make it really easy to get autoscaling groups of nodes with specific hardware.
What you want to run, and how, where, under what conditions, etc. is all specified in declarative files. Like I said, stateful services are not a problem anymore. K8S can deploy a database service as instances exclusive to specific nodes with the appropriate storage and networking config.
It's like an invisible assistant to automate what you would be doing yourself, or using several other tools to do anyway. A well-tested way to handle most of the common infrastructure issues for any deployed software in a single platform.
EDIT: I'll add that running K8S itself can be challenging, but that's because complexity can't be erased but only abstracted. I recommend using managed K8S to offload this.
They give you raw access to storage without redundancy. Which is great if your workload on top (like https://k8ssandra.io) already is redundant. Whilst making sure workloads stick to where their data is stored
Being able to get essentially the same result, but with our familiar processes, tooling, infra, that was really helpful getting things off the ground quickly and keeping them stable. We picked k8s/GKE way before that as a platform where a small set of abstractions allow us to do a lot of different things, and that makes it easy to debug unfamiliar workloads, keeping operations overhead low, and it worked out really well, and much easier for people with little ops experience to learn and get to a point where they could take care of their workloads themselves.
And it never felt we had to mangle k8s for our stateful workloads at all. Statefulsets, storage, node affinity etc. feel very much like organic building blocks and work really really well, with very few surprises. I guess k8s has come a long way as far as stateful workloads are concerned.
Whatever you use to manage the nodes directly is basically what K8S provides (and more) already, so you're really just replicating effort and complexity instead.
k8s orchestration is also useful for running multiple replicas of a database and using leader election to promote a secondary to a primary if the primary goes down. Doing this without k8s would require some kind of clustering anyway, to coordinate the leader election. So if you're already using k8s for everything else, you might as well also use it for this instead of introducing something else like pacemaker.
The one disadvantage is that leader election under k8s does need to be implemented from scratch, because it isn't a first-party concept in k8s. All it gives you is the first-party implementation of compare-and-swap for resources. So you have to make a StatefulSet for your replicas, make a ConfigMap to store the leader election data (which is where the CAS ensures atomicity), and write a sidecar container to handle the healthchecks and leader election and dynamic scale-up/down of replicas. You'll probably end up making an operator to do all that for you so that you can deploy a custom resource for your replicas and they're converted to StatefulSets with Pods containing the sidecars, etc automatically.
Of course, for the big popular DBs, someone has probably already written such operators for you.
I haven't yet seen good documentation on how these K8s databases handle upgrades. Additionally, they all seem very quick to stream from master to create a new node. That is great, for very, very small db's. But a 1TB db would be royal pain to do that with.
I've faced so many issues with the CSI driver which doesn't attach/deattach properly when you modify anything on the Statetfulset.
Additionally, you need to pin your DBs ("statefulset workloads") to certain Nodes based on the AZs since a block storage like an EBS cannot really be attached from Zone 1 to Zone 2 automatically.
Based on these limitations, I dont really see any utility/usefulness in running a DB on K8s. I get all of the other niceties that K8s offers, but here it makes the Ops work more complex than desired.
For example, use a storage service (like Rook/openEBS/portworx/etc) that uses node-attached disks to create a data layer. Then you can assign volumes to your individual workloads created and managed by the storage service.
This provides reliable and portable storage for your other services without having to tying them to lower-level details.
Having Postgres replicate the data across AZs or run Ceph on top of EBS to do it at the storage level are options. The complexity of setup is made simpler by running their respective operators.
There used to be bugs with this years ago but all storage providers now get this right.
I’ve not had issues with EBS/AWS, and affinity configs meant it was really easy to spread nodes throughout AZ’s.
On Azure though, volume mounts/unmounts were an absolute dumpster fire.
People say this but my first reaction to "let's put our DB in the Kubernetes cluster" is still to take two steps backwards and prepare to run in the other direction. It really goes against all of my better instincts (and unfortunately hands-on experiences).
Now almost all databases are run on VMs
In this case, you should be asking “will we all be running databases in clustered mode in the future?”
I think that answer still up for debate.
Sure, and containers are just plain processes on a VM. The exact code that controls processes in a VM, also controls processes inside a container.
As long as the underlying data volume is reliable, it really doesn't matter whether you're running it outside of a container or not. Volumes for K8s have a come a long way since its early days.
But running a database in K8s is implying (and it’s how the article sets it up) as a cluster. This is a big difference from running a solo instance on bare metal, in a VM, or in a container. That’s the point I was making to the parent comment.
To go up a level (comment), running Postgres in K8s doesn’t really get you much that you can’t also get from a VM (I’m thinking mainly about migration and setup). After that, the benefits come from running a cluster, which K8s can make easier to orchestrate over bare metal. But that’s a different architecture entirely.
and the answer is we will definitely be running DBs in clustered mode in the future. its not really a debate its an eventuality.
without clusters you lose on reliability and maintenance QOL improvements. without clusters its very hard to upgrade systems effectively leaving systems to stagnate and decay both in performance (Hard/Software improvements aka kernel updates or inbternal DB improvements) and reliability.
Or will they be on something more purpose-built, either designed from the start to handle storage more flexibly at a higher level? Or conversely something more specific and lower-level to extract maximum performance for their particular storage patterns?
And now almost all databases (meaning running instances, not software artifacts) are way slower than necessary and lie about durability.
now, that's a bit of a stretch, no?
I think this is a good ops team mantra. PostgreSQL unless you have a very good reason otherwise, and k8s unless you have a very good reason otherwise.
Decisions like this allow for economies of scale with tooling, monitoring/observability, etc.
Saying no to snowflakes is an important part of having a good ops team.
A persistent connection is antithetical to the design of kubernetes.
Persistent storage is hard to get right in kubernetes and puts a lot of pressure on the cluster being perfect, which is something a good ops team will tell you is hard.
If your application is stateless http services then kubernetes is great, but there are common things that don’t fit well, and forcing people into a solution which may require a complete rethinking is not very pragmatic.
https://cloud.google.com/config-connector/docs/reference/res...
https://aws-controllers-k8s.github.io/community/reference/rd...
A small database might be just $5 per month if its a thin slice of a big shared Kubernetes pie.
Or it could be $500K per month if it's a huge set of replicated server storing petabytes, replicated for dev, test, and prod.
I still think it's best to allow K8s to coordinate ephermeral containers with little to no persistent storage and let the DBs be run more traditionally.
Shameless plug: if you want to start running your own PostgreSQL with SSL, SELinux, automatic system updates, attached storage, and more... I have a simple demo as part of https://deploymentfromscratch.com/.
If I understand how it works, allocating memory to a container does not allocate memory for caching of data, and any IO done by any container can eject data from the cache, which means any IO intensive process will step on any other.
Zalando seems to fit the Zalando use case (not surprisingly), which is having multiple teams self-service their postgres clusters. They have a nice Web ui that is good for people who are not familiar with k8 to provision databases. However a lot of those features are a bit intrusive when you want to start a small project. The "teams" attributes are required in the clusters, for example, although you're not using them.
I ended up using Cruncy for my personal project because its a bit simpler, looked like having a bit more activity on Github, and slightly better documentation from my perspective.
Usually if it's a stateless app then the solution is using a global load balancer or other networking at the edge to steer traffic to the responsive regions. Databases also tend to have their own replication and distribution strategies handled internally that should be used, along with the above mentioned networking.
these manifests can be targeted with a Service.
However my preferred solution for databases inside k8s is still zalando-operator with spilo and local storage. it's rock solid. I once used kubedb and lost the database because of bugs, so I'm staying away from them. Crunchy looks solid aswell but it looks like it's a little bit more polished but harder to configure than zalando's solution which is more built to their needs.
You're suggesting that next time they should demonstrate Postgres on Kubernetes without the distracting Kubernetes part?