Even this blogpost doesn’t explain what and why’s of Kubernetes.
I have a docker container. I deploy them on vm. I use load balancer to split traffic. Could you please walk me through what problem Kubernetes would solve here?
Even this blogpost doesn’t explain what and why’s of Kubernetes.
I have a docker container. I deploy them on vm. I use load balancer to split traffic. Could you please walk me through what problem Kubernetes would solve here?
It does all of this by allowing you to specify a service architecture in configuration files and then actively ensures that this configuration is maintained even as the underlying state of containers change. You can specify things such as the minimum number of backend containers needed to provide a service, scaling parameters to add more backend containers as load increases, and you can tag nodes with different attributes so that containers are distributed and maintained with the appropriate amount of resources.
Kubernetes also provides automating various aspects of networking such as provisioning and configuring load balancers for service ingress as needed. It provides an internal DNS service which automatically registers names for deployments so that linked services can just refer to each other by name without any additional configuration. It can also manage things like SSL certificates which can be shared across multiple services.
Lastly, it provides you with a single place where you can store secrets and configuration values that these services require and again, you wire all of this up with configuration files which can be stored in a git (or other VCS).
But what if I told you that you needed to run that docker container on 5,000 machines with 2,500 load balancers in regions across the world? Are you going to SSH into thousands of boxes manually and run docker commands? Are you going to try setting up some monster ansible inventory to do the same? In practice the best minds in distributed computing have found those kind of practices break down at large scale--you just cannot reason or deal with individual machines when there are thousands of them.
This is where kubernetes comes in--it's an abstraction that lets you declare "here's the state I want, X machines running Y containers, all linked through Z services" and kubernetes will make it happen, period. It will take care of contacting thousands of machines, controlling the running containers, ensuring they stay running, handling failures, monitoring, load balancing, etc. You no longer think about problems in terms of low-level machines and instances, you think about the higher level objective like deploying code.
The beautiful thing is that it scales down nicely. A simple 50 line YAML file that declares running your docker container and load-balancing it with a service can easily deploy just to your local machine, or be scaled up to run on 5,000 machines by just changing a variable in the deployment scale. The same simple one-liner kubectl command kicks of either deployment and helps you monitor its progress. If you've ever worked in distributed systems it is really incredible to see this in action at scale.
What happens when your VM dies? Kubernetes would automatically bring it back up. It has health checks, and knows when containers die/crash.
What happens when you need another docker container due to traffic? Again, Kubernetes fixes situations like this. Kubernetes has a lot of built in support around scaling etc.
Also, what if your docker container doesn't need a whole VM? Say you've got 5 different docker containers (which all scale independently), and lets say 3 VMs. Kubernetes will distribute them across those VMs based on there resource needs.
There's a lot more, but that's kind of what I think of when you say 'Orchestration of containers'.
If you start a deployment ("your container", roughly speaking), the age of that deployment will keep counting up even if the container exits-- k8s restarts it of course-- and even if the k8s scheduler goes down-- since we want to be able to restart the scheduler without unnecessary service restarts.
At home, it's all on one computer in my basement, so when I reboot that box, the k8s reports come back and keep telling me the deployment has been there for xx days (just with some availability hiccups).
Even with a reboot, most OCR implementations (docker daemon, eg) will keep a pod around for 10 minutes until they reschedule it, so you'd see a restart count for the pod but probably not even a different age. It's tunable but IIRC 10 minutes is the default.
If you run your docker app on kubernetes, you get a lot of things for free with the platform (rolling no-downtime deployments, service discovery, auto scaling) that you’d have to set up manually if you were running your service on (say) EC2 instances instead.
It can be a headache to learn sometimes, but ultimately saves a lot of effort if your use case fits!
For better or for worse, you need something outside the Dockerfile that can run it. That can be you, if you want to type out 'docker run -p8080:80' etc. You could probably script it, but does your script do restarts, failover, etc?
- If you wanted to scale up or down the number of container processes running on your VM, you'd need to write some code that looks at system utilization. Autoscaling k8s clusters do that for you. They can even provision additional VMs ("nodes") for you during times of heavy traffic, or scale down to save money.
- Updating your app requires either logging into each VM, or writing an Ansible playbook to do that for you. By the time you've written a zero-downtime, health-check-honoring, contextually-aware Ansible playbook, you've made your own container orchestration solution.
- If you run multiple containers that need to talk to each other, you'd need to handle their networking. K8s gives you tools for handling networking between containers in the same namespace that allows them to communicate without exposing them to the wider internet.
- The ecosystem of utilities is as good (and sometimes better) than you'd experience in your VMs setup. cert-manager makes certificate management almost as easy as LetsEncrypt does on a single machine. Prometheus and Grafana are excellent logging and monitoring solutions (and, IMO, much easier to setup on K8s than ELK is within a distributed VM setup). Cillium provides extremely powerful and useful networking and security policies that leverage eBPF
- Changes you make to the configuration of your server won't carry over if you ever need to switch hosting providers, or (more often the case for me) just want to start fresh.
It's absolutely a huge learning curve, but eventually the complexity (mostly) goes away, and you're left with a reproducible method for deploying apps. So in the same way a rails/django developer might use an overpowered solution for their blog API, or a React developer may build a custom frontend when wordpress would also do... someone who's taken the time to familiarize themselves with K8s might find the familiarity and consistency of the interface enjoyable, even if it is clearly killing a fly with a sledgehammer.
Said another way, if Linux (or whatever) is the OS for your server / VM / host level / network device, k8s is the OS for your cloud application.
And, when k8s is implemented properly, it takes a lot of headaches that can come from dealing with the myriad problems that arise when your architecture goes beyond a basic handful of "tiers."
So it solves the bin-packing problem automatically rather than having to manually map out an efficient way to use your infra. It doesn’t reduce the complexity per say - the interactions can get complicated if you’re using stuff like node taints/tolerations instead of the more computationally simple “App X needs this much RAM but beyond that I don’t care where it lives”
If I find a bunch of raspberry PIs in my basement and want to have them join my fleet, I just do that and even if my fleet varies from 256 core CPU boxes with huge raid arrays all the way down to tiny raspberry PIs, the scheduling just works. Note here the broader pattern of abstracting away the physical hardware; it’s a really important concept to grok.
But that is what it means. K8s, ECS, even docker swarm are ways to orchestrate containers to do something useful.
Take your lb example. What happens when one of the containers you deployed or VMs you deployed to dies? How does it get restarted? Where does the lb send traffic?
If I add more physical hosts or scale the pods, k8s does everything for me like moving them around across the available resources
Deployments, scaling, logging, etc are some of the patterns it provides and the consistency matters. How many of us have worked at companies where deploying two services have been completely different? One team runs the jenkins pipeline while another team FTPS the files over. Now multiply that by several services and several tasks (logging, scaling, etc).
The benefit is in the patterns and standards it provides.
Seems to me like it's helpful to do an oil change yourself before you take it to the dealer, then understand what they do before you take it to jiffy lube the next time (or vice versa). You keep abstracting it away until you just give your credit card. :)
I use Kubernetes because it makes my application, its configuration, and its operations portable.
That's it.
A startup that grows to have hundreds of developers might transition from running managed VMs to "the cloud". One team sets up the network (virtual networks and some subnets).
As new employees join, they have no reason to interact with those teams who are effectively "hidden" so they deploy their stuff and perhaps wrangle with subnets and what not. Someone tells you that you need to attach subnet-a20w88vhuh4fuih to your resource and it will magically be accessible in the office.
Nothing is in charge of VM sizing so you've got people blowing hundreds or thousands on massive VMs when they're only using 10% of it and vice versa, teams whose application is choking but they don't really have a good mental model of say general purpose VMs vs memory optimised so they just bump up the SKU instead of being more efficient. This is happening everywhere as the company accelerates more and more.
It gets worse when you have a shared cluster say; for an entire team that is globally distributed and the new intern application is doing some weird O(n)6 computations and absolutely blowing the side out of every other resource you've got.
Now at this point, it's effectively a communication/culture problem but Kubernetes can "fix" some of these issues in a sense.
Network for the most part becomes abstracted away and what you're left with is defining security (what ports and protocols should I expect) on an application level, rather than on a security group level. It's kinda neat because these rules are localised to your application whereas they might have been configured manually in a cloud portal or via some terraform config owned by some team in the shadows.
Each of your deployed applications become their own isolated units called pods. A pod could be one or more containers but it's effectively a standalone slice of an application (ie the web frontend while a redis instance might be another pod). There are bigger abstractions to group application pods together but that's besides the point.
These pods get deployed to a cluster (a bunch of VMs) and cough "orchestrated" but the value here is that your containers might be running right next to some containers for the business team or the machine learning team and you would never know. You don't need to know either. The value, as foreshadowed above, is that if you're being a noisy neighbour, your container will either get rebalanced somewhere else or just shut down for exceeding memory usage.
I'm a bit flakey on this point but since each node in a cluster is a massive VM, there's no need to worry about over or underspending based on your computational use as well. You define the amount of memory you want to allow and you get matched to a relevant node based on how much capacity is available. As you gain more users, you just add more nodes. Before that, you might have been "reserving" say X thousand compute hours of certain VM SKUs or whatever. You might still do that but you could feasibly just have whatever your node sizes pre purchased making capacity planning pretty straight forward.
Generally, there'll be some team whose purpose is to manage said cluster so in a funny way, it somewhat revives the whole dev/ops split in that your compute team generally know the nitty gritty of networking and what not while your developers just deploy an application and it "lives on Kubes".
I may have a missed a bunch of stuff but hopefully this outlines some of the more "people" issues a bit? It's half and half useful but also it can be used as a technical fix to a social issue.