Disclosure: I work at Google on Kubernetes.
Disclosure: I work at Google on Kubernetes.
I've been wary of running anything stateful on K8s for other reasons than the stable identity stuff. Historically, Docker's I/O has not been particularly robust — well, Docker has not been particularly robust. You also have to be very careful about resource limits and requests.
In particular, you have to ensure that the QoS class "Guaranteed" [1] kicks in, because the last thing you want is for Kubernetes to think that PostgreSQL is something that is happily rescheduled.
[1] https://github.com/kubernetes/community/blob/master/contribu...
tl;dr: do you want to be a DBA? do you want to be a cluster admin? If you want to be both you can run databases on Kubernetes. If you don't want to be a DBA, use a managed service.
Non-containerized environments (VMs, bare metal) are well-understood and fairly predictable; Docker and Kubernetes are moving targets. There are "known unknowns".
As an example, I am involved in a project that uses continuous delivery for deploys, and we recently had an incident where the GKE node ran out of inodes because Docker was keeping all the layers of dead containers around. Kubernetes handles this correctly: Best-effort pods were evicted and the node went back to being ready once the inode use went down. PostgreSQL would have stayed up in this case, but it was still a little surprising that GKE, which is nominally managed, does not have a process in place for cleaning up Docker's garbage. (Edit: Apparently it's a bug [1].)
When discussing Docker and Kubernetes, there are very often edge cases that come up. For example, until a couple of months ago, Kubernetes had a serious bug that would often cause the wrong network volume to be mounted to a new pod. You can be the world's best containerized-database admin and still be surprised by things like this, which have nothing to do with the nature of containers, and everything to do with the specifics of the software that controls them.
"containerd is all we wanted from @docker in @kubernetesio and none of what we didn't need: kudos to the team!"
For example, I'd be able to upgrade by just starting another container and replicating into it, then routing traffic over. I'd retain the option of rolling back by routing traffic back if anything goes wrong. I'd be able to spin up/down redundant read-only replicas simply by changing the number of replicas. And other apps would be able to piggyback on Kubernetes' discovery system (typically you let service names resolve into IPs/SRV records through DNS). And so on.
You can do anything you want with classical VMs, of course, it's just more static and slower: Provisioning another VM, configuring it (Puppet, Salt, Chef, etc.), connecting persistent storage, reconfiguring everything to point to the right places, configuring DNS, etc.
In short: If you use Kubernetes as your orchestrator, why would you not want to use it for everything?
But it just seems like you're trying to shove a very stateful management problem (the schema, the associated storage) into a system designed for mostly stateless stuff. And that way lies a giant hassle.
Plus we're certainly not at huge scale, but db rollbacks are nontrivial, and we have to be somewhat careful about schema modifications.
Thanks for your comment!
There's nothing about containers that says they need to be stateless, it's just more convenient that way because you can treat every container as expendable, redundant, fungible commodities. But the real world needs state, and Kubernetes is capable of managing it. In particular, the pattern pretty much required for stateful apps on Kubernetes is that of using an exclusive-access network block device. Since only one container (or pod) can mount it at any given time, you can ensure that the state follows the container around.
It may just be our use case -- we have enough data (low TB) that cloning is not a fast thing to do.
Thanks again!
We deploy a small cluster (1 master, 6 nodes) at our startup that started misbehaving last week. All of a sudden three of the nodes went down - one became unresponsive and two had the error "container runtime is down." We couldn't ssh into the unresponsive one, but according to AWS the machine was fine, still receiving network requests and using CPU.
Since we couldn't diagnose the issue, we spun up an entirely new cluster using kops, but started seeing the exact same behavior later that night, and again over the weekend. Three nodes were in a not ready state, for the same reasons (unresponsive and container runtime is down). Right now our only solution to solve this issue is to manually terminate the EC2 instances and rely on the Auto-Scaling Group to create new ones. In the mean time, Kubernetes tells us that it can't schedule all of our desired pods, so half of our jobs aren't running, obviously an undesirable situation.
A handful of questions I have about the situation: Why are these nodes going down? What causes a node to go unresponsive? Why does the container runtime go down on a node and why doesn't it get restarted? Why doesn't Kubernetes destroy these nodes when they've been out of commission for 3-4 hours?
Any help would be appreciated!!! I've been looking through half a dozen log files and gotten zero answers.
IIRC we improved garbage collection settings in the latest kops (1.5.1), so if you were running out of disk, using the latest kops should fix everything. It's also easy to reconfigure to use a bigger root disk if you're churning through containers faster than GC can keep up. But if it's something else we can try to diagnose it as well!
> Why doesn't Kubernetes destroy these nodes when they've been out of commission for 3-4 hours?
We should, I believe. I actually thought we had an issue for this very problem, though I can't find it. I'll open a new one if I can't track it down. There is maybe an argument that we should fix the root cause, but there's an unlimited number of things that can go wrong, so we need to do both.
(edit: Gave up on finding the existing issue and opened https://github.com/kubernetes/kops/issues/2002 )
I rerolled with 100G and I've seen zero problems since.
Kubernetes isn't responsible for the lifecycle of its nodes. It can run in a DC where "destroying a node" might mean paging a tech to turn off a server. Something external - in your case, kops & your ASG - is responsible for the nodes that Kubernetes runs on. That's a deliberate design choice.
It should make a correct decision not to schedule work there, which it sounds like it did.
Given that, your other questions are hard to answer. kubelet is a process that runs on the nodes. So is docker. If you can't get into the machine to diagnose the fault, I'd encourage you to set up some monitoring/log shipping off the node so you can see what the state was when it failed.
There's nothing inherently "Kubernetes" about this diagnosis - it's more EC2, node/kernel/OS and Docker troubleshooting, in that order.
If you can't get to the machine, there are a million reasons why this would be the case - but ssh is a totally separate process, it's way outside of Kubernetes. VERY commonly, you've run out of memory and processes are fighting among themselves (especially since EVERYTHING seems to be failing), but this is total speculation. OS issues are common too - I've spun up clusters switching from one distro to another, same config, and everything worked great.
Disclosure: I work at Google on Kubernetes.
The Container Optimised OS is what GKE uses on Google Cloud Platform https://cloud.google.com/container-optimized-os/docs/
It's conceptually very similar to CoreOS' Container Linux, so I might try that if I were looking at Kubernetes elsewhere and wanted a container-only OS.
If I am running an environment with multiple purposes - some container hosts, some regular machines - I'd err on the side of "who is my current vendor/what does my ops team support and know best".