Treat Kubernetes clusters as cattle, not pets
zitadel.ch
zitadel.ch
Sorry, but I feel like I landed in crazy-land. Kubernetes is already an exercise in how many layers you can insert with nobody understanding the whole picture. Ostensibly, it's so that you can isolate those fucking jobs so that different teams can run different tasks in the same cluster without interfering with each other. Hence namespaces, services, resource requirements, port translation, autoscalers, and all those yaml files.
It boggles my mind that people look at something like Kubernetes and decide "You know what? We need more layers. On top of this."
lol... what?
A CEO of a 20,000+ person company literally sees projects as revenue sources and a deadline, nothing more.
To get it delivered, he/she walks down the chain of managers until they gets answers. It's might as well be a black hole.
And it's a red flag. Means that IT is running the show, not engineering.
Heh. Service Mesh!
(...sorry =_=)
https://github.com/kubernetes-sigs/cluster-api-provider-nest...
it’s k8s clusters all the way down
> It boggles my mind that people look at something like Kubernetes and decide "You know what? We need more layers. On top of this."
I'm not involved with this project at all, but yes, I look at the massive Kubernetes deployments I've worked with, and conclude that toil and various other kinds of problems could be reduced with higher-level abstractions for declaratively managing configuration for all of these clusters.
If you wanted to run workloads on 100k servers, how would you do it? Would you have a team manually configure each cluster individually, or would you consider that there are a lot of similarities between clusters, and you only want differences that are chosen intentionally, and look into some abstraction to help keep the complexity down?
If 100k isn't enough for you, is there any number of nodes at which you'd consider building some tools support for your work?
What about data center count and change rate? There are a lot of business models that benefit from running workloads closer to customers, so it's useful to have deployments in many datacenters and clouds across the world, and for this to grow over time. Do you really not see why some people are interested in being able to bring up a cluster in a new DC by updating a few configs rather than starting from nothing by hand with kubeadm init?
I apologize if I've misrepresented your position by misunderstanding something. I feel like I'm missing or misunderstanding something from your perspective, but I don't know what it is.
More generally, I think it's extremely normal for people to build higher levels of abstraction as things grow. It's usually worth considering abstraction or architectural changes in a system every time it grows by about 10x.
Some of the most painful platforms I've supported in my career have been those that didn't do this, and instead tried to grow through brute force. Something can be fine to do once, troublesome to do 10 times, and overwhelming to do 100 times.
For myself, I think it depends on how leaky the abstractions are. The leakier, the fewer layers the house of cards can stand. That is, it's more about where you slice it than how many times.
As you say, there's still a fundamental amount of complexity involved in operating a system, and the best abstractions limit the amount you need to care about in a given context.
Something I like about declarative configuration specifically is how well it helps you move information from ephemeral human memories and habits into something more-reliable.
Most places I've worked usually had some weird ephemeral things, oral history about what needs to be treated differently, many systems that needed conversations to understand, etc. The more you limit your use of global mutable state (interactively changing things in production), the less room there is for important stuff to live in people's heads.
When I can spin up on a new service just by learning how a tool works and reading their configs, then I have all of the "what" and "how" and "where" there to look up at any time in a consistent way. There's much less room for weird state or needing to do rituals or broken staircases or whatever. When I talk with people, I can focus much more on the "why" questions.
I might be more sensitive to this than some people, because I've got a shitty memory, but the more I'm able to just check a thing the better. Theoretically documentation can fill a lot of the same role, but I've never worked anywhere with consistently up-to-date and comprehensive documentation for what's deployed. When it's the only way to deploy at all, it's no longer optional.
It feels like the same sort of thing as a good static type system to me. You can do local reasoning with just what you see, and limit unexpected action at a distance. The types are mandatory, and machine-checked, so you can't get them wrong in certain ways, and you can't just skip writing them like you do sometimes for unit tests that would cover similar verification otherwise.
But... Most don't run 100k servers. Most don't run 10k servers. I was working with a very large consumer insurance company recently and while they are really large for their segment, they're well under 15k servers in production. And containers aren't much of a thing. There's a lot of overengineered architecture going into not-so-hyperscaler/FAANG production environments compared to what Kubernetes' original use case was designed to solve for. The common consequence appears to be increased complexity and less operational system knowledge .
You're right that most people don't run 100k servers. In fact, most people don't run any servers at all, and don't use kubernetes. Some people do, and some of those people use approaches like this.
The comment I was responding to, by my reading, seems to express incredulity that anyone would do this, and seems to imply that this is a bad, counterproductive idea that nobody should ever implement. The most notable quote is "I feel like I landed in crazy-land". Do you read it differently from me?
My intent in my reply was to describe some of the use-cases where this is helpful. It's not only FAANG that run more than a small handful of clusters; there are plenty of smaller businesses that need to run large numbers of clusters.
I don't know anything about the company that wrote this article, but I found https://zitadel.ch/usecases on their website, and given that description, it sounds quite plausible that they run many small clusters in many clouds and datacenters. They probably don't need a ton of compute, but I wouldn't be surprised if they had latency needs for being close to their customers, or other specialized requirements that motivate dedicated clusters.
I also don't know anything about the insurance industry, so I'm curious to hear if I've guessed incorrectly, but it doesn't seem like the kind of business with especially high compute needs. I believe you when you say that your company has no use for this.
Can you believe me when I say that there are genuine problems these systems are trying to address? I keep hearing about companies who fund their employees building unnecessary pointless over-engineered production automation, but I have yet to ever encounter one in real life. Every SRE job I've had so far has been on a team that really wanted to invest time in improving infrastructure and automation, but couldn't get time away from the toil for it. If anyone can recommend a company that over-invests in infrastructure automation and architecture instead of under-invests, I'd dearly love to try it and see if it's as bad as I've heard.
Even without high compute needs or large numbers of clusters, there's still a lot of benefits you can get when you can move things to declarative configurations and away from global mutable state and human minds. I agree that this can go wrong, but that doesn't mean it's categorically worthless to try.
On the other hand, if you're just looking to gripe about it having gone wrong for you, could you share some war stories? I may be feeling idealistic from frustration with low investment in tools support and production automation lately, so maybe I could use some horror stories of it going badly to scare me straight. :)
In contrast, i think that Docker Swarm does a much better job at smaller scales, because:
- it uses way less resources than Kubernetes (which matters on smaller nodes)
- it is included in the install of Docker and therefore is easy to launch
- it supports the Compose format, which is often used for development with Docker Compose
- it is relatively simple, yet handles most of the concerns of smaller scale projects (i.e. excluding things like autoscaling or CRDs)
- there are solutions to manage it with web interfaces as well, like Portainer, if necessary ( https://www.portainer.io/ )
Or, if you need Kubernetes in a non-managed format, i'd suggest that anyone take a look at projects like K3s ( https://k3s.io/ ), which attempt to cut out all of the unnecessary components and features of Kubernetes, that most people won't use.In my opinion, going with Kubernetes as a whole or even full Kubernetes distros (or building your own clusters for prod) is akin to choosing Apache Kafka because you want to build an event driven system, and then needing to throw a team at the thing to actually keep it running smoothly, instead of just choosing something simpler, like RabbitMQ or ZeroMQ.
I'd argue that most companies out there don't even have "SRE jobs", but instead it's a duty that gets thrown upon regular devs who also need to deliver features. It's probably easy to overestimate how capable most companies out there are in regards to throwing resources at problems to make them disappear, mostly due to a lack of said resources.
I look at "if you've got 100K servers..." and ask "why does anyone needs 100K servers?", rather than "how would I manage that?".
Each layer adds complexity and reduces performance. Less layers are better. Write Unikernel applications
I think the point is so the teams can run services at all. Isolation isn’t the point, it’s breaking a dependency on an ops team (Kubernetes in a sense is ops as a service).
I hear people say Google runs their business this way. So perhaps it makes sense that Kubernetes came from them.
We're already there. Cue "We invest in teams not ideas"[1,2] and "We're pivoting since we couldn't get a product-market match, but have millions in the bank"[3]
1.https://www.forbes.com/sites/ilyapozin/2015/10/08/4-reasons-...
2. https://www.gsb.stanford.edu/insights/do-funders-care-more-a...
This statement is less about adding another layer. It's more about not falling in to the same trap that we did with VMs in 2010--managing them by hand, producing unreproducible configurations, "tending" to them. The same pattern is happening with Kubernetes, but it fundamentally runs on computers at the end of the day so it's prone to similar types of failures.
They'll do warehouse scale computing with borg operating large clusters. borg is at the bottom.
The workloads spanning dev, test, and prod then run on these clusters. By having large clusters with lots of things running on them they get high utilization of the hardware and need less hardware.
It's amusing to see k8s used in such a different way and one that often uses a lot more hardware while driving up costs. Concepts Google used to lower the cost.
Or, maybe I read the papers and book wrong.
I like the idea of higher utilization and better efficiency because it uses less resources which is more green.
No, that's exactly how it works. You have clusters spanning a datacenter failure domain (~= an AZ), and everything from prod to dev workloads runs on there, with low priority batch jobs bringing up the resource utilization to a sensible level.
You can do the same thing with k8s, you just have to trust its multitenancy support. You have RBAC, priority, quotas, preemption, pod security policies, network policies... Use them. You can even force some workloads to use gVisor or separate prod and dev workloads on different worker machines.
Kubernetes can be upgraded. I've watched nicely done upgrades happening 4 or 5 years ago. I've watched simple upgrades happen more recently. It's not unheard of. Even in public clouds I've upgrade Kubernetes through many minor versions without issue.
I would argue it's more work to create more clusters. You need to migrate workloads and anything pointing to them. It also would cost more as you have to run more hardware.
Of course it is the cluster of Theseus since every single bit of compute has been entirely replaced :)
You can also throw in some smaller clusters that replicate the 'main' clusters' software stack and have some workloads whose downtime does not impact production jobs (not just dev workloads, but replicas of prod workloads!). These can be the earliest in your rollout pipeline and serve as an early stage canary.
I think an important thing to know about K8s is Omega failed to replace Borg, and then the Omega people created K8s. So K8s does not necessarily descend from Borg, and not all of Borg's desirable attributes made it into K8s.
Borglet/borgmaster releases definitely break workloads. I recall one where something was rolled out to all the test clusters, passed all the tests (except one of ours) and was about to be promoted to increasing percents of the prod clusters. Whilst debugging why our test (which is not part of the feature rollout) broke we realize that if this had rolled out to prod, it would have broken all Tensorflow jobs, and would have been a major OMG.
But yeah, most of the time, the release process for borglet and borgmaster is fairly fast and fairly reliable.
A big part of this is that outside Google, the number of people who have to operate k8s as a fraction of all k8s users is way higher than the fraction of borg users who have to operate borg, so there's a lot of stuff in k8s that is 'end user experience' comforts and affordances for operators.
Cluster API takes the above and can in-place upgrade clusters for you. It's pretty awesome to see first hand. Bumping the Kubernetes version string and machine image reference can upgrade and even swap distros of a running cluster with no down-time.
My hope is that projects like k3s will manage to cover that small scale spectrum of the market.
One attraction of k3s is you can get a "real" production-grade K8s running on a single machine, or a small cluster, very quickly and easily. That can be great for learning, development and testing, certainly.
K3s is actually targeted at "edge" scenarios where you aren't running in a cloud, and don't have the ability and/or desire to dynamically adjust the cluster size.
You can scale a k3s cluster easily enough manually, adding or deleting nodes, but to get cluster autoscaling you'd need to do environment-dependent work to make it automatic.
If you do want that, then things will likely be quite a bit easier with a managed cluster like Google's GKE, or even possibly an self-installed cluster using a tool like kops, kubespray etc. (It's a while since I've used any of those, so not sure which is best.) AWS EKS is another choice, although its setup is a bit more complex and I wouldn't really recommend it to someone getting started.
And for production scenarios, integration with cloud load balancers, network environment etc. is similarly going to be easier with a provider-managed cluster. It's all possible to do with k3s, but it's more work and a steeper learning curve.
https://research.google/pubs/pub43438/
https://sre.google/sre-book/production-environment/
are the best public documentation entrypoints that I know of.
To be fair, google runs its own datacenters with teams doing research & optimisations on all levels of the stack: hardware and software and most importantly an amount of engineering resources converging to infinity.
The rest are stuck with VMs that share network interfaces and have to monitor CPU steal, understand complex pricing models, etc. Engineering resources are scarce so most companies will over-provision just to be safe and because profiling the application and fixing that API call that takes too long is expensive, will spin-up another 50 pods.
This analogy really bothers me. Cattle are expensive. They are an investment. You don't put down an investment just because it got sick.
If you have a sick cow you will in-fact call your local large animal vet to come and treat it.
However, the idea that whenever anything weird happens you should just kill you server/cluster and move on without doing any sort of investigation seems like a recipe for disaster. That's a great way to mask bugs that may in-fact be systemic in nature, where they are or will eventually cause service degradation for your customers.
I would hate to work in an environment where bugs are ignored and worked around instead of understood and fixed.
Ideally that goes into an async queue of issues though and someone finds the root cause and that goes into a queue of issues to fix, which is actually burned down.
I suspect what is happening more often is that the whole stack has so many levels and the SREs responsible for it all don't have the visibility into the stack they need to debug it all, so they use their SLOs as a club to ignore issues as long as they're meeting their metrics until it becomes a firefighting drill.
A pile of cargo culted best practices and SLOs replacing hands on debugging.
I'm far more likely to ignore and work around a bug instead of doing a proper investigation when I've got pressure to get production back up because this server/cluster is a Special Snowflake that must be fixed in-place.
Hardware fails, and bugs happen. There's no getting around it. Automation to handle this case is a good part of any strategy for identifying, understanding, and fixing bugs.
There’s nothing remotely disgusting about making sausages, but people say “Laws are like sausages - it’s better not to see them being made.”
I stopped correcting them long ago.
In my opinion it needs to be re-engineered completely into a super slim product that is not tied to all these crazy things.
You can spin up a local multi-node cluster using kind[0] in 1 minute on 6+ year old hardware. I know it's not the same but I really have to imagine there's ways to speed this up on the cloud. I haven't spun up a cluster on DO or Linode in a while, does anyone know how long it takes there for a comparison?
My Terraform scripts get a HA K3s cluster in Google Cloud VMs in less than 9 minutes, which in my opinion is fantastic.
I've worked on KIND and on clusters on clouds (at Google, but on multiple clouds) and both can be very quick, if anything there's still low hanging fruit to make KIND faster that I'd expect a production service with more staffing to handle.
KIND is Kubernetes, typically on much weaker hardware :-)
Within a few minutes is a perfectly reasonable expectation even for "real"-er clusters, see e.g. under 4 minutes:
- https://kubedex.com/google-gke-vs-azure-aks-automation-and-r...
- https://stackoverflow.com/questions/53839292/is-a-3-minute-g...
An Ingress or a service of type LoadBalancer will create a load balancer in AWS that's tied to your cluster, but that's the whole point of Kubernetes, it'll spin up the equivalent resources in Azure or GCP or DigitalOcean.
And I can't blame them.
The real problem they suffered is actually that Kubernetes isn't fundamentally designed for multi-tenancy. Instead, you're forced to make separate clusters to isolate different domains. Google themselves run multiple Borg clusters to isolate different domains, so it's natural that Kubernetes end up with a similar design.
[0]: https://jzelinskie.com/posts/youre-not-running-vanilla-kuber...
Disclosure: I worked as an engineer and product manager on CoreOS Tectonic, the (now defunct) Kubernetes used in the post.
Disclaimer: I am working with ORBOS
Invitation accepted! :) I've been dying to see an ALTS-like [1] thing that works with Kubernetes. I really should be able to talk encrypted and authenticated gRPC-to-gRPC without ever having to set up secrets or manually provision certificates, dammit.
[1] - https://cloud.google.com/security/encryption-in-transit/appl...
It's fairly subjective what types of isolation define "multi-tenancy" which is why there hasn't been progress made despite SIG and WG efforts in the past. While you do not believe network isolation should be included, there are plenty of developers working on OpenShift Online which may disagree. OSO lets anyone on the internet sign up for free and instantly be given their own namespace on a shared cluster full of untrusted actors.
It worked out great for us -- upgrading Kubernetes was easy and testable, never worried about code drift, etc.
[0] https://blog.asana.com/2021/02/kubernetes-at-asana/ (Not written by me).
I looked it up and for the cost of putting them up in the “hotel” I could have them euthanized and buy new cats three times over.
I would never do such a thing, but I did use it to guilt them into not complaining about the hotel.
It didn’t work very well.
Even if you have full bare metal automation and thousands of machines that seems like unnecessary waste.
It can add up to millions per year if you aren't auditing your costs.
* Orchestration of jobs (restart, start); this can be achieved easily without the complexity of k8s
* Sidecar loading; Literally the easiest thing to do with normal VMs.
and...?
Let's not forget the potential security implications of not keeping things properly isolated.
People who were around when provisioning on bare-metal was still a thing already learned all these lessons. Somehow it seems they have been forgotten by all the people driving hype around Kubernetes.
I’m super not hypey about kubernetes, mostly because the complexity surrounding networking is opaque and built on a foundation of sand... But let’s not argue things that aren’t true.
So... why are these issues open?
https://github.com/kubernetes/enhancements/issues/1029 https://github.com/kubernetes/enhancements/issues/361 https://github.com/kubernetes/kubernetes/issues/54384
> additionally: if your node becomes unhealthy then the workloads would be rescheduled on another node.
Well of course, but you're going to run into that issue (likely) on all of the nodes where the offending service lives.
> But let’s not argue things that aren’t true.
If what I've said is untrue, looking at open GitHub issues and the Kubernetes documentation is certainly no indication. That's a massive problem all by itself.
The second is a tracking issue for a KEP that has been implemented but is still in alpha/beta. This will be closed when all the related features are stable. There's also some discussion about related functionality that might be added as part of this KEP/design.
The third issue is about integrating Docker storage quotas with Kubernetes ephemeral quotas - ie., translating ephemeral storage limits into disk quotas (which would result in -ENOSPC to workloads), vs. the standard kubelet implementation which just kills/evicts workloads that run past their limit.
I agree these are difficult to understand if you're not familiar with the k8s development/design process. I also had to spend a few minutes on each one of them to understand what the actual state of the issues is. However, they're in a development issue tracker, and the end-user k8s documentation clearly states that Ephemeral Storage requests/limits works, how it works, and what its limitations are: https://kubernetes.io/docs/concepts/configuration/manage-res...
That's just not the case.
Ephemeral storage has resource requests/limits in pods.
Traffic shaping/limiting can be accomplished using kubernetes.io/{ingress,egress}-bandwidth annotations. It's not as nice as resources (because there's not quotas and capacity planning, and it's generally very simplistic) but you can still easily build on this.
Pods can also have priorities and higher priority workloads can and will preempt lower priority workloads.
> Let's not forget the potential security implications of not keeping things properly isolated.
For hardware isolation isolation, you can use gVisor or even Kata containers.
> People who were around when provisioning on bare-metal was still a thing already learned all these lessons. Somehow it seems they have been forgotten by all the people driving hype around Kubernetes.
Kubernetes explicitly aims to solve resource isolation. It was built by people who have decades of experience solving this exact problem in production, on bare metal, at scale. Effectively, Kubernetes resource isolation is one of the best solutions out there to easily, predictably and strongly isolate resource between workloads _and_ maximize utilization at the same time.
High overhead can be automated away, google ORBOS.
Also if a cow dies, people don't just buy a new one. It represents quite a loss of profit. Also represents a big potential problem on the farm that people will want to resolve - they're you're money makers, if they're dying it's an issue.
Losing a server was often devastating, even with backups. "Don't run your builds right now, people are visiting the website" also sounds a lot like "Don't connect to the internet right now, i need to use the phone".
I mean there's "correct" ways to implement the whole pet servers concept but .. why? It's fun, sure, but it's also a waste of time, and when you want to be productive it just tends to get in your way.
Truly, if your software team headcount is under 500 why are you running k8s?
A well implemented mini cluster can and will save you so much time in later maintenance and deployments.
Don't even get me started on monitoring.. k8s murders everything else here
Oh god! Please treat your cattle better!
Is it really though? I for one am glad I didnt jump on the bandwagon early. A lot of the articles popping up nowadays mentioning the downsides of GitOps make a lot of sense.
Costs are generally less of a concern, but having one way of running, operating, and writing services allowed our dev team to move faster, share knowledge, etc.
One API as abstraction with shared processes eases the pain for the people relying on a platform.
Everyone seems to be thinking they're Google or Facebook, or both.