It's an incredible mess of artificial complexity. You finally understood kubernetes and can work with the manifest files? Oh, no, no, no, nobody does that. You need to at least learn helm and a dozen of other kubernetes ecosystem tools! Every company has a different setup for it.
* I have a GKE cluster running to learn Kubernetes. Once you understood the basics it's indeed a simple way to get applications up and running, but my points stand. I am far away from enterprise readiness.
To add to that, I think there's a very specific way you must do things to get K8s "right" and reap benefits. Or maybe it's a question of scale. There were honestly times when I thought maybe we could've stuck with good old servers. When I left the team was trying to redo microservices---this time properly---by introducing event queues and pubsub and finally split that database.
Love that workplace but I'm honestly glad I didn't get to participate in that transition. Sounds like "artificial complexity" as you put it. Would love to hear if they pulled it off though or not, a project postmortem basically.
The practical results for me were:
* Having to write a CLI app to reduce the time wasted and spit out simple config templates for our actual (very much not Google-scale) use case
* Whenever something didn't work, throw it over the wall and hope the DevOps gods deemed us worthy
I like tools that have sensible defaults and try to stay out of your way. K8s fails on both counts in my experience.
Why is this surprising? You always need people to manage stuff. Was there any point where Kubernetes was advertised as a way to NOT have sys admins/devops?
Don't even get me started on "serverless"....
Kubernetes is made for the people that have to run your systems and keep them operational (ops/devops). It makes our lives a billion times easier on two fronts. Disclaimer, I am referring to a managed solution like GKE, I can imagine running k8s on your own hardware would be harder.
Uptime & Stability: autoscaling, health checks, replacing nodes, restarting containers. These are all things that an Ops person does not have to worry about during day to day operations.
Visibility: This is the biggest thing for me. People hate CLIs but kubectl and it's standard commands (get, describe, etc) make my life so much easier as it allows me to quickly figure out why something is breaking. A typical troubleshooting session goes like:
- What customer are you looking at? (kubectl get namespaces) - What service is failing (kubectl get pods -> kubectl describe pod <pod name> -> kubectl logs -f <pod name> -c <container>
and the error is often quite obvious.
This visibility is powerful and is something you can easily treat your developers to do on their own. Even though Ops owns the systems they do not have to be a black box to the Developers that use them.
My organization is playing with EKS and while perhaps it is different from GKE I feel like I have minimal visibility into if the control plane is healthy. I've had several instances where all of my node groups have become unhealthy for various reasons and EKS failed to spin up new nodes to resolve the issue.
Observability starts with logging and metrics. K8S already puts out lots of events on what's happening which you can use to debug, and there are many 3rd-party monitoring solutions if you need more detail.
If you don't have this with EKS then use something like LogDNA and NetData to see the logs and metrics yourself.
Its literaly the best option currently for what it is doing.
It has an unprecedented support behind it as well.
Multiply Vendors support it through a certificate k8s managed service.
It solves really problems out ouf the box like load balancing, ingress, cert management, autoscaling, health checks, autorepair.
It allows for simple IaaC.
Is it young? Yes. Do we need more people with more expierence? yes.
Is this a problem? No.
> Is it young? Yes. Do we need more people with more expierence? yes.
Of course this is a problem! It's costly (in time, money, and security) to pay your team to ramp up on Kubernetes.
The question is, what is the actual benefit? 99% of companies don't need it at all.
Its costly and risky to run 100 VMs, maintaining them and keeping them up to date, monitoring them and knowing when they are no longer needed.
It is cost ineffective to have security audits on 100 VMs, maintaining access to them, auditing whats happening on them.
It is a ton easier to allow someone only access to one namespace, limited ingress domains (which get provisioned automatically) and allow them to only run non root containersl
The question is not what the actual benefit is (its clear and i mentioned it in my paragraphs before), the question is will kubernetes ever become so lightweight, stable, easy to use so that its feasable for normal people to run it on a 3-5 node cluster in small companies.
And please lets ignore all those small companies where people log into their VMs by hand and maintaining them by using snapshots and cloning existing VMs. Those small companies might and should just migrate to cloud managed services completly.
A tiny fraction of companies need to run 100+ VMs. It takes a lot of traffic to require that scale, unless you're just throwing money at infra to avoid optimizing your code (which can be cost effective).
A large chunk of companies should just be on something like SquareSpace or Shopify, another large chunk should be using off-the-shelf services like Azure's App Containers or Amazon's Elastic Beanstalk/Lightsail/whatever else they have now, and some internal services can easily run on a single VM with no orchestration at all.
All that stuff is really expensive at scale, but unless you're a large company, your engineering time is going to cost way more than just using whatever managed services your cloud offers.
> ...maintaining them and keeping them up to date, monitoring them and knowing when they are no longer needed.
This is not the only alternative to using Kubernetes. You can automate all of this without Kubernetes, or you can use a cloud provider's managed Kubernetes service. A lot of companies get by just fine with Heroku.
If you assume people are using practices from 2005 (or that cloud providers don't already provide a layer of abstraction on top of Kubernetes), of course Kubernetes looks better.
I did not say that everyone should migrate to kubernetes as far as i rmemeber but kubernetes to me is not a fad and it fixes real issues companies have.
This is what will get lost on hn: 99.99% of companies don't need FAANG-level capabilities.
I bet that 80% of people moving to K8s just follow hype and don't even benefit in any netto positive way.
And it's probably 99% but I'm just too afraid to be wrong.
Here's an example.
Blog post says: "After transitioning each service, we enjoyed many benefits of using Kubernetes in production, including much faster and safer deploys of the application, scaling, and more efficient resource allocation"
Wrong. Actual analysis by the engineers of one of their migrated jobs (urgent-other sidekiq shard) says their cost went from $290/month when running in VMs to $700/month when running in Kubernetes. They tried to use auto-scaling but failed completely and ended up disabling it:
https://gitlab.com/gitlab-com/gl-infra/delivery/-/issues/920...
Kubernetes looks like a massive LOSS in the case of this service. The perceived scaling benefits didn't materialise at all, and their costs more than doubled. They also spent a lot of engineering time optimising startup time of this job: 7 PRs and complex writeups/testing were required. Just so they could try to auto-scale a job to reduce the hit of the hugely multiplied base costs. If not for Kubernetes that eng time might have been spent adding features to the product instead.
It solves really problems out out the box like load balancing, ingress, cert management, autoscaling, health checks, autorepair
Most of these problems are created by Kubernetes, like "ingress" which is a term only Kubernetes and low level networking uses, "cert management" which often isn't required in more traditional setups beyond provisioning web servers with the SSL files, "health checks" are a feature of any load balancer and the whole point of VMs is to avoid the need to repair hardware. Finally the demand auto-scaling as seen in this case is (a) possible without Kubernetes and (b) not actually working well anyway.
Frankly this set of bugs, writeups, PRs etc is quite scary. I worked with Borg at Google and saw how it was both a force multiplier but also a huge timesink in many ways. It made some complex things simple but also made some simple things absurdly complex. At one point I had a task to just serve a static website on Borg, it turned into a three month nightmare because the infrastructure was so fragile and Google-specific. A bunch of Linux VMs running Apache would have been far faster to set up and run.
Note this is just one service out of many and has very low latency requirements, and due to our slow pod start times in this case it didn't make sense to auto-scale. We are auto-scaling for other services in K8s.
> says their cost went from $290/month when running in VMs to $700/month
These numbers represent a snapshot in time where we are overprovisioning, for this service for the migration. You are correct that cost benefit here remains to be seen for this service. We are doing things like isolating services into separate node pools which may not allow us to be as efficient.
Safer and faster deploys was a huge win for us, for this service and others. This of course is compared to our existing configuration using VMs and Chef to manage them.
disclaimer: blog post author
This is only one of the shards of one service that we migrated, and we migrated many more.
I did a high level writeup of this and many other things we've observed in https://about.gitlab.com/handbook/engineering/infrastructure...
if you are curious. I am happy to have a conversation with anyone about our experiences in more detail, not to convince anyone that k8s is the "one-true-way" (because I am still not convinced myself), but of the benefits this type of change can bring.
> They tried to use auto-scaling but failed completely and ended up disabling it
That issue is just one of many though. We disabled it at the time, but we have it enabled for a number of other services.
Regardless of how I personally feel about K8s, I have to say that the migration we are doing for GitLab.com is generating set of benefits that goes far beyond just moving to a new platform.
One of the largest benefits I've seen so far is that it was a great forcing function to resolve some long running architectural challenges, and is making us think more about how the application can run at a very large scale, without being at the very large scale. Things that we could get away before, like the issue you referenced there, we can't anymore.
Disclaimer: I am one of the people involved in this migration.
I mean you mentioned yourself you worked at google right?
Perhaps you just haven't actually experienced the issues kubernetes is solving?
How often have you seen that certificates expired? I have seen that. Plenty of times. Its not an issue creating that lets encrypt cronjob, its still something you need to do right.
Security Updates? Have you seen how many companies run with old non updated VMs?
Disk full due to logs? Yes seeing this regularly.
Memory leak on a service and someone needs to restart it manually until someone else fixes the issue? Yes!
Requesting a VM, hardware, infrastructure, getting it and the whole lifecycle management of it in the backend? Its real.
Ansible Scripts, puppet or just bash scripts and a word document to tell you how this magic machine was set up? Yepp.
Kubernetes solves all those problems.
Your static website on borg, if it sill runs, has probably still a valid certificate, is running on a secure infrastructure, is equally configured on every instance and not weird on 1 of 6 servers and just runs.
A smart person taking responsibility for all of this, costs you much more then just a few hundred bucks a month. And you need that person. With Kubernetes, this person now can manage and operate much more servers under his/her fingertips better easier and more secure then if it would have been vms.
And in my personal experience: That shit runs more stable because that shit can restart and being recreated and it just solves a handfull of shitty memory or disk full issues.
Borg is/was great for running huge numbers of services at truly massive scale, when those services were developed entirely in house and done in the exact way it wanted services to be done. It was a terrible cost the moment you wanted to run anything third party or which wasn't written in that exact Google way, and the costs were especially high if you didn't need to handle huge traffic or data sizes.
Borg and Kubernetes don't auto-magically do sysadmin work. A ton of people work in infrastructure at Google. Software updates still need to be applied etc. In Kubernetes that means rebuilding Docker images (which in reality doesn't happen as the tools for this are poor, so you just have a lot of downlevel Ubuntu images floating around).
And worse, the whole K8s/Docker paradigm is totally backwards incompatible. At Google this didn't matter because the software stack evolved in parallel with the cluster management. Programs there expected local disk to be volatile, expected to be killed at a moment's notice, expected datacenters to appear and disappear like moles. They were written that way from scratch. But that came with a terrible price: it was basically impossible to use ordinary open source apps. You could import libraries and (carefully!) incorporate them into Borg-ized projects, but that was about it.
When this tech was reimplemented and thrown over the wall to the community, the cultural expectations that came with it didn't come with it. So I've seen situations like, "whoops, we deleted our companies private key because it was stored to local disk and then the Docker container was shut down, help!". This was even reported as a bug in the software! No, the bug is that your computer is randomly deleting critical files for no good reason, and normal software does not expect that to happen. How about the way in a Dockerfile you have to write 'apt-get update && apt-get upgrade' if you want a secure OS image when the container is rebuilt? If you put the two commands on separate lines it will appear to be working, right up until you start getting weird errors about missing files from Debian mirrors, because Docker assumes every single command you run is a pure function! And then we get to security.
Screwups seem to follow Kubernetes/Docker around like flies. The tech is complex and violates basic expectations programs have about how POSIX works. When it goes wrong, it leads to mistakes in production.
Now there have been a few trends over time:
1. Hardware has got a lot more powerful. It has been outpacing growth in the internet and economy. Many more businesses fit in a smaller number of machines than when Borg was designed (an era where 4-core systems were considered high end).
2. Cheap colo providers have been driving the cost of powerful VMs down to the ground.
3. Linux has got easier to administer.
These days setting up a bunch of Linux VMs that self-upgrade, run some services via systemd etc isn't difficult, and properly tuned such a setup should be able to serve a monster amount of traffic (watch out though for Azure, which seems to overcommit capacity pretty drastically and their VMs have very unstable performance).
Most businesses that are deploying Kubernetes today quite simply do not need a million machines. Even companies that give away complex services like GitLab, as we can see from this thread, they don't really need huge scalability. It's just nice to imagine that the business will experience explosive growth and that growth is now automated, but ironically, the effort to automate business infrastructure scaling takes away from the sort of efforts that actually grow the business.
As for the other benefits, you can write a systemd unit that is much simpler than Kubernetes configuration that will give you auto-restart, including if memory limits are hit, you can view the activity of a cluster easily using plain old SSH or something like Cockpit, it will handle log rotation for you out of the box, sandboxing likewise, and so on. And of course the unattended-upgrades package has existed for a while.
I agree you need someone to do admin work, whatever path you choose. Having used Kubernetes, and the system it is based on, and plain old Linux, my intuition is that the base cost of Kubernetes is too high for almost all its users. Too many ways to screw it up, too many ways for it to go wrong, too much time spent screwing around, and too expensive. If you become another Google or Facebook then sure, go for it. Otherwise, better avoided.
But you reach quickly enough an size where a central managed kubernetes cluster is very versatile for your whole company.
1-2 People take care of the cluster, the other teams then use that kubernetes cluster for your build, hosting of staging envs etc.
And then you have a very small infra team which is much better able to provide those services to others internally much easier and safer.
Because there have been other 'fads' still being strong today.
Kubernetes solves real problems which have been hard for a long time. Its the first thing you, as an infrastructure team, want to have to be able to provide your teams a manageable environment for yourself.
mesos, nomad, docker swarm and co.
Developers don't want a VM and you can't manage and maintain VMs if someone else is doing something with them.
Cloud native is here to stay; It will affect and already affects tools, applications and architecture.
I want a VM because I want my dev environment to exactly replicate the production environment, down to the kernel. It also means that if I leave my employer I can just delete my VM on my personal machine.
We've also not had any issues with maintaining production VMs, our release pipeline packages up our code into a debian package and them creates a VM image from that. This is then deployed across all regions. If you need to roll back just deploy the previous VM image version. GCP handles killing old instances and starting new instances for you.
Feel free to be happy with your setup.
Also i'm not able to determine if your setup would still be much better (faster, easier to maintain etc.) if you would set completly on containers instead of a VM as i'm not aware of your workload at all.
Yes if you are developing C/C++/Assembly then Kubernetes/containers might not be for you.
But for several other companies (e.g. Java users) the kernel version is not important.
So we are back to same argument. Just because that you personally don't see value in something, doesn't mean that it is a fad.
Then maybe you agree that Kubernetes brought some advantages that until recently only Java devs enjoyed to the rest of the world (PHP, Python etc)?
Kubernetes is a kludge to solve problems that should be sorted out by the respective language communities.
Then there is the whole issue that interpreters without JIT/AOT should only be used for scripting.
GP is the one who made blanket statement 'developers don't want a VM', not me.
We've also not had any issues with breeding horses....If I just need to go somewhere I just ride my horse...
I don't see any value in cars(k8s). They are complex (oil changes, fuel consumption etc...)
Please replace with Ansible, Chef, Puppet, HP Vault, Solaris Zones, OpenMosix, Bewolf, Grid Computing, MTS, CORBA, JEE .....
An argument based on seeing hype come and go, sellig conferences, books, certifications, consulting, training, and naturally one selfs curriculum.
Speaking of which,
"Kubernetes Certified Application Developer (CKAD) with Tests"
https://www.udemy.com/course/certified-kubernetes-applicatio...
OpenMosix, Bewolf and Grid computing were mostly confined on research circles. I never saw them took off like K8s did.
So comparing K8s with OpenMosix is a bit unfair...
Eventually there will come the post-Kubernetes generation and the cycle does yet another turn.
Do you now think progress is bad?
Ansible and co did and still do a great job.
kubernetes is not just a thing, its also a paradigma shift.
It sounds like something new is just bad because its new?
My first containers were called HP-UX Vaults in 1999.
The cycle continues its reboot.
Its getting refined, standardized and widely used.
k8s for me is the shift to a declarative infrastructure. Easy to put in your source code repository.
Also as a goody you get industry leaders offering a certified k8s experience.
And all of it is opensource and free.
A fully platform independent orchestration platform.
I think that K8s is much too complicated for medium-scale workloads, so as soon as something makes it easier than K8s to scale and manage workloads at that scale, K8s will be out of that space.
It doesn't make much sense to migrate to k8s if you don't have an issue. And it doesn't make sense to wait for the next thing to happen when you have an issue you need to fix now.
Now you can have all this complexity even in other languages than Java.
If you're running microservices (to the point where you have so many small services that having an asg per service), than the above strategy does waste resources compared to k8s I guess... But you'd need a lot of microservices to justify the significant overhead k8s has.
For K8s you need to manage the vm's/os as well but that part is mostly trivial as it's highly focussed on immutable/disposable infrastructure.
The golden rule is not to pretend to work at a FANNG when not having the same problems.
Auto-scaling is something I'm suspicious of. I saw many experiments with that at Google when I worked there, under the name of elasticity. They found the same thing GitLab found here: it's really hard and the expected savings often don't materialise. Even in 2012 their elasticity projects were far more advanced than what Kubernetes or AWS provide.
Most cloud auto-scaling efforts appear to be driven by the high base costs charged by the cloud providers in the first place. I've seen a few cases where simply moving VMs to cheaper colo providers saves more money than auto-scaling VMs or using managed Kubernetes in AWS/Azure.
GCP does offer custom metrics, but I couldn't figure out how to tie them to the LB scaling logic.
Kubernetes solves a specific set of problems. Just because you personally don't have these problems, doesn't mean that Kubernetes is a fad.
Its features do not justify its complexity and the cost of operating it.
Then there was that tiny physics research center close to Geneva creating what was one of the first grid computing platforms.
But what do I know, that was almost 15 years ago.
Yup, pretty happy with our move from k8s to ECS. Almost there.
This 'k8s' fad is the open source version of what you are using.
Its like you would say 'monitoring is a fad' we migrated away from monitoring and are now using amazon CloudWatch and are super happy with it.
That would be AWS EKS. ECS is very different beast. Wins big in infra-as-a-code (CFN). And simpler, fully integrated, and you get full AWS support, not the case with EKS or hand rolled kube.
> .. monitoring is a fad ..
We did some pretty fancy things in monitoring. ELK and more. And yes, it uses some useful bits of CW.