Databases on Kubernetes is fundamentally same as a database on a VM
twitter.com
twitter.com
With a stable operator it's theoretically similar, but it is not the same, by any stretch, if only because I guarantee your incident rate will be higher when performing changes to the database, all other factors being similar. You now are also exposed to a bunch of other changes, the shenanigans of k8s networking, a missing daemonset, someone was tweaking istio that messed everything...
The argument of "don't run it on k8s" isn't because you can't run it physically. It's because all those other things are going to increase your failure rate, and those failures will be way more complex to solve and so your time to recovery will also go up.
You'll have to deal with this complexity already for the things that Kubernetes is going to give you actual benefits on, doing databases should be at the bottom of that list. Use your complexity budget wisely. If you're already at the bottom of the list, you look at your team and you go "we have extra capacity and are very confortable with the complexity of our current incidents", go ahead, likely it is the right time for you if you like adventure.
I understand what the saying means, but it gives people so many bad ideas about the value of theories and science or how to work with them. In practice you start out based on general and overly simplistic theories, and then expand the theories in the direction you need to solve your problems, and then that creates a new theory you base your practice around.
Wandering around willy-nilly without testing isn't good practice, while going deliberately with lots of testing and logic is what we mean when we talk about theory.
For practise to be condensed into theory requires predictive power, and not all differences between practise and a theory are predictable -- yet.
That said, I agree with your overall point. If one finds a difference between practise and theory, it's because one tried to use the wrong theory for the situation at hand.
You are correct, though, in that if you want to make these practices more useful and grow your overarching theory, they absolutely must be approached with lots of testing and logic. The trick is knowing where it's most important to do this & when.
Many areas of practice are left untouched by theory because there is no obvious compelling benefit.
Well said
Running an app with state (db, storage, etc) in like this is not that hard. We ran all of the low level storage on borg at google when I was there. It worked well.
[0] https://github.com/CrunchyData/postgres-operator/issues/3476
Another example: https://github.com/zalando/postgres-operator/issues with 445 open issues. Why?
Maybe I'm wrong and this is all a good sign of progress, but my impression is that the entire k8s ecosystem is held together with reused duct tape.
Sure, a lot of it is nice-to-haves but some are absolutely essential for even a basic postgres deployment even if you used VMs... Like, why is replication so difficult, why do i need all these extensions and why do i need to deploy a bouncer as a separate service in 2023?!
My opinion is colored by seeing HA postgres deployed on k8s for no good reason, then having to deal with the consequences.
It just seems like once you're in the k8s world, the likelihood that your system will be massively over-engineered goes way up. This isn't the fault of the operators, or k8s itself, it's more of a cultural problem I think. And resume driven development is real.
It's basically just a bunch of already well thought out open source tools like patroni, pgbackrest and postgresql combined into an operator. So it's not really magic.
And the backup to s3 feature in pgbackrest is so solid that I've used the clone cluster and restore cluster feature a few times, just because it's so convenient.
But this is in environments that serve only thousands of users, not at all comparable to huge startups, or environments where ms latency is important.
Can I just provide it storage from a k8s-host and it replicates it to the hosts where the other replicas live via software or do I need to provide it a StorageClass that supports distributed volumes?
This is the one thing I cannot wrap my head around and apparently there isn't that much information (read benchmarks) on the implications of using different kinds of storages for databases on k8s (or at least I can't find anything)
If configured like a daemonset, you are garanteed to always match disk to pod instance ordinal (postgres-0 get disk 0, postgres-1 get disk 1).
So basically a way to run multiple postgres instances.
But the devil is in the details.
I could try this out myself to be honest, but I didn't quite get to it until now
Which is linked to from https://hub.docker.com/r/crunchydata/crunchy-postgres
1: https://access.crunchydata.com/documentation/postgres-operat...
> By participating in the Program and accepting these terms, you represent that you understand that (i) absent an active Crunchy Data Support Subscription, the Crunchy Developer Software is unsupported and (ii) Crunchy Developer Software is intended for development purposes only, (iii) Crunchy Developer Software may not address known security vulnerabilities, and (iv) Crunchy Data is relying on your representation as a condition of our providing you access to the Crunchy Developer Software.
Whereas at their Github repo is there's an actual license (Apache) and says stuff like
> Grant of Copyright License. Subject to the terms and conditions of this License, each Contributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable copyright license to reproduce, prepare Derivative Works of, publicly display, publicly perform, sublicense, and distribute the Work and such Derivative Works in Source or Object form.
Contact crunchys email with your specific usecase and ask if it's within their free license if you don't believe this. expect a gigantic bill however as you're definitely in violation. (The bill will be their normal licensing fee, not because of the previous violation)
You are probably confusing PGO with their commercial certified offering.
Without checking the docs I would imagine you could do a little hack where you restore from pgbackrest into a clean postgres 14 cluster.
They have failover, cluster elections, node auto joining mechanisms.
Vanilla postgres have none of that.
Interesting, I have run my hobby projects by Nomad satisfiedly and are looking for ways to run serious workloads. Would you like to share more wisdom? How do you accomplish things above? Thanks.
I'm unsure if something like open policy agent can directly work with the orchestrator and may have to be at the application level.
https://www.hashicorp.com/blog/nomad-service-discovery
https://www.hashicorp.com/blog/nomad-1-4-adds-nomad-variable...
I definitely agree that databases don't benefit from running under containers - they are pet-like (and so don't benefit from fast spin up or massive horizontal scaling) and tend to require host-level tuning (which breaks the container abstraction).
What Kubernetes brings is a well-principled orchestration framework that can easily be extended with custom operators for workloads such as databases that need it. (In fairness, he does refer to this in passing at the end.)
K8s was always really crap for persistent storage. On AWS you can always dump to EFS or the managed lustre (what ever thats called). (which is better than attaching block storage.)
I suppose this is mainly a thought for projects like neondb/cockroachdb/stackgres (who I haven't heard of but was linked in the thread). It might be reasonable if you need incredibly many db instances, but for the general business who needs "a couple" of database instances, I can't imagine that putting Kubernetes on top would ever serve you better. I'm staying as far away as I can.
EDIT: part of actions, codespaces and packages are not run on k8s, but 80% of github services are
That's really not the endorsement you think it is.
For example “users cannot resume code spaces created before the incident” sounds a lot more like an application level problem.
Blanket statements like that should be taken with a grain of salt.
F500's are not one thing. You don't have to scratch deeply to find teams running production DB's on k8s (ignoring or accepting the trade-offs, of which there are many including working with vendors and existing DBA's and their solutions) and you'll find DBA's evaluating the same and other trade-offs for themselves.
I personally think that running DB's on multi-tenant k8s with nodes that weren't specifically allocated for it is strapping in for a bad ride.
StackGres is "just" a platform for running Postgres on Kubernetes. It helps you deploy and manage HA, connection pooling, monitoring, automated backups, upgrades and many other things. That you have a tiny Postgres instance; or hundreds of beefy clusters with many instances is up to you. It's not a distributed database (like the other ones mentioned), it is still "vanilla" Postgres.
Disclosure: Founder of OnGres (company behind StackGres)
I was very skeptical of data on Kubernetes when we first started, in part due to some initial experience with Kubernetes in 2018 but mostly due to prejudice against change. Overall it has worked out great. Here are 4 of many things we've learned.
1. Most modern databases are distributed systems. You don't just set up a single node but rather several or even dozens of nodes. Well-written operators make this relatively trivial even though it's quite complex underneath. In fact, the simplest way to learn how to set up a ClickHouse cluster is to bring it up under the operator and then look at the configuration on each container. That's how I learned it.
2. Kubernetes portability is overall quite good. We ported our cloud from AWS to GCP in 8 weeks. We've since expanded to run in many other environments as well.
3. We map ClickHouse server containers 1-to-1 to VMs spawned using Karpenter or native node groups. It makes it a lot easier to reason about performance, including things like network bandwidth to storage.
4. ClickHouse is still basically a shared nothing architecture where individual servers own patches of storage. Kubernetes enables a great scaling model if you use VMs attached to block storage--you can scale nodes from 2 to 64 vCPUs in a few minutes, plus you can easily extend volumes. This scaling model is in my opinion highy under-rated for databases. It's decoupled compute/storage that really works. With Kubernetes you get it essentially for free.
It's not all roses. Containers create new failure modes. You can't just ssh in, look at logs, and fix things. Pod crash loops [0] can be very problematic. Certain failure modes like bad EBS volumes (kinda alive states) are hard fix if your operator cannot quickly replace a node. And operator bugs create a new class of very-hard-to-debug problems. The best solution to all of these problems is not to have them, which means you need to focus--often for years--on operator reliabily and your day 2 infrastucture, such as monitoring.
[0] https://altinity.com/blog/fixing-the-dreaded-clickhouse-cras...
Disclaimar: I work for Altinity
Each database container is attached to a single disk volume and they are not fungible.
There are ways of moving beyong this model, but you have to innovate at the API layer.
For example Microsoft Service Fabric, has reliable collections and queues. Its a single Dictionary and queue accessible by your application that is partitioned and replicated with your application. It scales with your application and always writes to local disk.
Why I am unsure of merits of this implementation, I imagine we need something like that to trully have good approach to databases in the new paradigme
https://learn.microsoft.com/en-us/azure/service-fabric/servi...
I have to real experience with it, just found their approach interesting
I'm a fan of k8s, but the rational of running DB's on it has always seemed specious to me.
More like a trick poney, actually, because you'll have to teach it a lot about what to do with the several container lifetime hooks and states.
But hey, if you really need to run your postgres instances in kubernetes (which you don't because cloud postgres instances can be fully automated) kubernetes operators are like the world's expert in poney training working for you.
Your pet, your problem. one operator, world's problem.
What's the practical difference?
For example you can not have 2 postgres instances, each with their own locking mechanisms, check whether transactions are serializable, since they don’t know about the locks in the other instance.
This is very basic stuff so I wonder whether people who argue for treating databases like web servers lack basic training in this area.
The crux of my argument is: you probably don't need postgres in kubernetes on the cloud (just use the damn cloud api to manage your database), but if for any reason you want to, manage a postgres with kubernetes it's like managing it with systemd.
Way more involved, but doable. Great part is: just as vast majority of people don't write their own systemd units or postgres management scripts, they can use kubernetes operators and benefit from the open source knowledge and ecosystem.
I much prefer to be able to treat the database instances like cattle, while the conceptual "cluster" and its associated backups should be treated as a pet, as you said.
So you can run a db on kube ok, but a "production like" setup will be keeping other workload off the same kernel. So the kube bit is because it's what you prefer to operate, and you might as well be using vms.
Taint the node pool so that only the database can tolerate it.
Presto. Now when you scale up N instances, N new nodes are created.
The method of snapshotting uses a different plugin, but apart form that, its essentially the same.
I suggest listening to his Space in full
You should NOT run your databases on Containers. As Werner Vogels says, "eventually, everything will fail" at scale, and the more complexity you add to a system, the higher the chances. K8S is an unnecessary layer.
Plus, don't even get me started on the security aspects of this choice...
I suggest listening to his full space on the subject
Don't tweet something that implies a very obvious conclusion, and then say "well if you read the reams of fine print you know he doesn't mean that".
Did you read beyond the first tweet in the sequence?
His point is that the decision depends on what you are trying to do. Running prod database, bad idea. Letting developers or automation spin up ephemeral databases, good use case.
The idea that k8s and vms are the same is that neither handles day 2/3 issues for you and both methods require a lot of expertise. i.e. deploying on a VM doesn't solve the hard parts either
Exactly this.
[1] https://criu.org/Main_Page [2] https://www.youtube.com/watch?v=wCb1Rfoy7Fk
No, you run stuff on Linux, using kubernetes to manage it.
Is deploying a database binary and mounting
a volume on a VM really the hard part?
The nice thing about containers is that you are independent of the underlying infrastructure. Be it a VM or bare metal.I can easily spin up a Docker container on my laptop. Takes a second or so.
Spinning up VirtualBox or something is so much more hassle.
It's true for non-trivial cases, and exceptionally false for some things. Like databases. A lot of the sysctl tuning for large database environments has to be done on the host, either directly, allowing unsafe sysctls and restarting kubelet, running with privileges, etc.
This is not "independent of the underlying infrastructure" or "the VM and my laptop have the same kernel tunables set". They don't. And it matters.
Most applications don't.
And even if the application I work on does - then only in production. Not on my laptop. So on my laptop I will simply chose a docker image that resembles production as closesly as possible. I will not use a VM.
You’re probably right there’s a magnitude more tiny apps but who cares. A large number of us still are interested in and need to solve for non-trivial cases.
None of these are solved by "optimize kernel-parameters to the max".
Only when you don't need to store data for any length of time. The point of containers is that they don't store state. You have to manage that.
And for some cases you can get away with it. But for most websites, keeping a single container up for many years isn't a great strategy. (Yes, I know there are ways around this. but they are not as simple as "docker run postgres")
Vagrant used to be a thing and it would spin vms up pretty quick.
There are differences technically, between docker and vms, certainly, but conceptually they are the same, aren’t they?
There’s no real
I have not used a container today. Let's use one:
$ time docker run --rm debian:11-slim echo hello
hello
real 0m0.528s
user 0m0.015s
sys 0m0.023sAs a random example, Azure full-clones 127GB disk images by default. It takes over a minute to create a VM. Booting form a cold start is sluggish because there is no sharing with other tenants, hence no caching.
Using the same hypervisor (Hyper-V) I can clone out a Windows server VM and boot it in about 3 seconds by simply using delta cloning. Subsequent boots of it or any of its siblings is just over a second!
Containers and Kubernetes are throwing the gloves down and will force the competition to pick up the pace.
AWS EC2s start within a couple of tens of seconds max. AWS Lambda, Google Cloud Functions, Google Cloud Run start within ms (Google Cloud Functions used to have a cold start problem where sometimes they'd take up to a few seconds to start, but that has been fixed).
But overall, I'd say VM platforms are somewhat a thing of the past. Nobody cares about running an OS, what you need is the things inside (your application, database, etc.) so fundamentally a VM is an abstraction at the wrong layer, a means to an end. Don't get me wrong, they're still here and aren't going anywhere, but should no longer be the go-to outside of a few specific cases - bare metal, containers and "serverless" (running on bare metal) is where it's at.
The difference is my laptop doesn't provide any guarantees of stability, uptime or throughput.
installing a DB is as simple as "apt install $db"
> The nice thing about containers is that you are independent of the underlying infrastructure
I mean you're really not. You're dependent on the machine, the OS, and the kernel version. Docker on windows is a VM, from what I remember. I assume OSX's docker is still VM based too.
Otherwise the operator concept makes it even easier. Zalando operator for PostgreSQL takes care of ha and backup
Yeah, but its not.