Experiences with running PostgreSQL on Kubernetes
gravitational.com
gravitational.com
Also I'd be curious to see if "most people" actually use Ceph vs network storage like EBS volumes where AWS guarantees me that I won't have data corruption issues in exchange for money.
Disclosure: I work at the company that published this post.
I read it differently (albeit, I have much more context). I read it as a cautionary tail that Kubernetes makes it easier to get in trouble if you don't know what you are doing - so you better have a deep knowledge of you stateful workloads and technology. Perhaps obvious to some, but still a good reminder when dealing with a well-hyped technology like Kubernetes.
They will likely get better at it over time. Until then, either run it outside of K8S or deal with the warts.
Our experience building and running 24x7x365 HA Kubernetes clusters for clients is almost exclusively in air-gapped or on-premise environments where EBS (or even an IOPS QoS guarantee) does not exist. Most SaaS and managed service provider teams we encounter would much prefer to leverage a cloud provider's primitives than DIY.
If you want any semblance of DB performance from Pg as a container, you've got to give it its own data volumes, mount that into its container, etc. Not sure what the benefit of containerizing it would be. Shared tenancy on kube workers makes your performance unpredictable as well, and if you use affinity to run it on a dedicated worker, what's the point?
I think his comment about stateful sets is more focused about running data stores on distributed filesystems, or network attached storage. This isn't a good idea generally, and I am not going to go into why, but Cassandra and Etcd advise against this for a reason, its well documented. It wouldn't just be using a 'better' storage system, it would be designing and implementing a very complex operator as he mentions.
The real take away which I think we should all realise is that, this is shiny tech to run mission critical data stores on, unless you know what your doing and want to invest a lot of time and effort in to it. Right now it's probably better to use RDS/Cloud SQL, or just run a traditional postgres setup with a SAN and decent failover that is tried and tested.
I missed that part. Did they mention it in the article? Perhaps they should mention that kind of stuff more prominently. Numbers... and such.
Source? I've never heard that claim, and in fact we've lost data on EBS (though it has been a couple years)
You could always sync to another volume in another region if you were concerned about a longer term region outage. Note that RDS already mostly uses EBS.
I generally avoid abstractions for persistence layers, but I believe I'm the minority and I believe I'm going against what the industry desires me to do.
I'm not convinced this is a good idea /yet/. I did, however, see some interesting docker/kubernetes integration with ScaleIO (clustered filesystem) which cut out huge chunks of the disk I/O pipeline (for performance) and was highly resilient.
The demo I saw was using postgresql, the dude yanked the cord out of the host running the postgresql pod.
Quite impressive in my opinion.
https://github.com/thecodeteam/rexray
https://github.com/kubernetes/examples/blob/master/staging/v...
"Your scientists were so preoccupied with whether or not they could, they didn’t stop to think if they should."
Moreover, the use case he's discussing is quite unusual: Gravitational takes a snapshot of an existing Kubernetes cluster (including all applications inside of it) and gives you a single-file installer you can sell into on-premise private environments, basically it's InstallShield + live updating for cloud software. So, running everything on Kubernetes opens entirely new markets for a SaaS company to sell to.
This level of zero-effort application introspection hasn't been possible prior to Kubernetes, so that's another reason to use it for everything: it promises true infrastructure independence (i.e. developers do not have to even touch AWS APIs) that actually works.
Sounds like a nightmare to me.
Postgres on Kubernetes seems close to the DevOps ideal for making lots of little "edge databases" and limited backends for uninteresting webservices...
It's absolutely fine to pin your Postgres to a handful of pre-selected hosts with locally attached storage fully exposed to a Postgres container. Yes, this won't be semi-magically "moving databases around" (not needed in most cases) but you'll still be getting other k8s benefits I listed above.
But even if you feel adventurous and want to have a fully dynamic storage under your RDBMS, there are tools for this now in open source / commercially supported form [1].
You need to override Kubernetes' built-in controllers to customize them to the details of the database, for example when is it safe to failover, or when is it safe to scale down. Outside of Kubernetes these decisions are made by cloud services like RDS, or manually by people with knowledge like DBAs.
To put that knowledge into Kubernetes you need Custom Resource Definitions (CRDs).
I know know if it is any good, but I just found a project called KubeDB that has CRDs for various DBs: https://kubedb.com/ and https://github.com/kubedb
A big part of the obstacle is that preserving state in a distributed environment is just hard, no matter what your technology, and the failure cases are generally catastrophic (lose all data, everywhere). This is true both for the new distributed databases, and for the retrofits of the older databases. So building DBMSes which can be flawlessly deployed by junior admins on random Kubernetes clusters requires a lot of plumbing and hundreds of test cases, which are hard to construct if you don't have a $big budget for cloud time in order to test things like multi-zone netsplits and other Chaos Monkey operations.
Making distributed databases simple and reliable is a lot like writing software for airplanes, but clearly that's possible, it's just hard and will take a while.
If you have async replication with a large amount of lag, well of course your gonna lose data if master goes down if your not careful. Regardless if your using kubernetes or not...
Can anyone explain why these failure modes are kubernetes specific? They just sound like things you have to think about regardless if your running a HA cluster...
What I took away from TFA is two-fold: clustered-DB administration is hard to automate (e.g., there's a reason DBAs exist), and a lot of the tooling DBAs use now has to be rebuilt to function in a k8 environment, or replaced.
The first struck me as blindingly obvious; the second I find highly relevant, personally.
This is why I'm excited about CockroachDB; it hopefully will make operating DB cluster easier.
Most of the issues presented in the article relate to data loss during network segregation and leader election. Those are important, but distributed systems are generally a bit more explicit in their CAP compromises.
It's not postgres specific either. It will happen whenever you have a cluster or HA solution which is not intimately cognizant of the app it is supporting.
Shit, it is hard enough to get any database extremely reliable even with their own solutions (Oracle RAC, Postgres SLONY, mysql replication).
Adding Kubernetes / Docker to that mix makes for an interesting life. Caution is advised.
If you have N servers that are deployed by hand or, better say, with some automation ran by hand - you can be sure things are not moving at random. If you schedule a payload in K8s (or Docker Swarm or whatever) - the scheduler will only ask you for preferences, but will otherwise make decisions on its own. Like deciding to drain a node because disk is getting full or whatever. Or just schedule an updated service to another node because it felt like doing so. And most of software - Postgres included - doesn't have any idea about all of this.
Some of the points raised (like leader election) are not specific to orchestration, but some (like persistent storage for containers) are sort of unique to containerized world. Unless there is a crazy sysadmin who can walk into server room and randomly shuffle drives. :)
At work, we run OpenStack on Kubernetes with Postgres for persistence, and it's entirely okay if Postgres fails for, say, an hour, because we don't have a defined SLA on the OpenStack API. The important thing is that the customer payloads (VMs, SDN assets, LBs) keep working when OpenStack is down, which they do.
There are two main things to understand. First is that your application is running in a distributed environment and will encounter data loss if it is not designed for it. Second, even if it is properly designed to run in a distributed environment it's also has to be aware that it's running on kubernetes, configured specifically for it.
The data itself is on a self-replicating storage (and Postgres obtains the proper exclusive locks when using it), so I don't care if it runs on Kubernetes or a Raspberry Pi.
I could have just said that even knowing nothing much beyond cursory reading about kubernetes, the article comes across as a excercise in attaining disappointment thru hurried assumptions about what constitutes both a silver bullet and the daemon to be dispatched from unruliness.
The part that is disconcerting is the introduction to the article as a interview with the CTO, but it only takes a turn for the worse almost immediately by admitting to I'm production deployment of the solution, to which the subsequent admission to encountered difficulties is not compounding the sin do much as burying this entire excercise beneath condemnation, if I simply put down the impression conveyed. This has to be at the very least terrible PR. I'm increasingly concerned too, about the abundance of misapprehension of not only the capabilities of file systems but just fundamental design constraints, at s level of understanding that I would have expected to be fired for from a operations position in any of my customers. Have I missed the redeeming features in my haste to comment? It just feels so imbalanced and insecure to be so forthright about the level of accomplishment that's claimed.
Ideally StatefulSet does help a bit. Such as with StatefulSet you have DNS and hostname like service0.namespace etc. And you can change StatefulManifest without updating the pod. Pod only get updates when we deleted it(OnDelete updating policy) or rolling update with staging parition which mimick real server behaviour where we can pause/restart process on a server.
However, what I realized is resource scheduling and EBS volume.
1. Soon I realized the node run db pod should only run db and I make a dedicated node pool for it.
When this occurs, It feels like eventually I'm provisioning a server to run this workload.
2. EBS volume cannot mount to other zone. So it really annoying when I kill a pod and it cannot start because the volume cannot attach to node in other zone.
And when we need to upgrade Kubernetes itself in an immutable way, mean kill old node and bring up new node. It's a pain to control that process carefully to avoid casscade node re-balancing/replication.
More over, the ability to easily goes in server and edit/tweak config is lost with K8S. We have to use ConfigMap, some trick of init container, entry point script to generate custom configuration file etc.
An example is broker id in case of kafka or slave-id in case of MySQL. In other words, I feel like running stateful service on K8S is no longer a joy.
Once I move these stateful service out of K8S, suddenly everything is so smooth. Running and upgrading K8S itself become a walk in the part.
Also, on AWS, when you have a large amount of server, the chance that you got AWS notification about node replacement/rebooting (old hardware, host migration) is very frequent. Dealing with these when all of node have stateful service running on is not easy.
This doesn't seem to me like a reason not to run PG on K8s?
Would something like Mesos resource reservation mechanism with persistent volumes do the job? When you run on premise you usually want to recover from a temporary failure or reboot and maybe run a special admin script if you feel like the node is not going to come back anytime soon.
1. https://github.com/sorintlab/stolon/blob/master/doc/faq.md#h...
this was on about 33rd minute of the presentation
He suggested , that if you need to use , synchronous replication, to use PostgreSQL 10 with synchronous quorum replication, so that it will not block transactions if one of the replicas fail.
It is a tradeoff though. With synch rep, you are at a minimum adding the cost of two network trips to the latency of each write (distributed consistent databases like Cassandra pay this cost as well, which is why they tend to have relatively high write latency). It turns out that a lot of users are willing to lose a few transactions in a combination failure case instead of having each write take three times as long.
Postgres also has some "in between" modes because write transactions can be individually marked as synch or asynch, so less critical writes can be faster. I believe that Cassandra has something similar.
No thanks.
One admin is more then enough to cover and master FreeBSD Jails and many other technologies ... no just Kubernetes.