As a generic rule I'm against using complex tech where there's no need, but I feel like we're way past that on k8s.
As a generic rule I'm against using complex tech where there's no need, but I feel like we're way past that on k8s.
k8s gets shit for one very simple reason: it forces teams to deal with issues that they're not necessarily ready to deal with just yet. I've seen it happen over and over and over.
Small startups are able to get away with not using auto-scaling, having bad secret management, not running minimally sized containers, etc, etc. Are they going to solve these problems eventually? Yes, if they grow and don't die. But until they grow they simply can't afford dedicating a lot of effort setting up Hashicorp Vault with proper secret rotation (just as an example).
One of the companies I worked at was a 10 person startup that was trying to find product market fit at a break-neck pace. We had customers, big ones. It was a non-trivial piece of software too. We ran the whole thing on a very beefy EC2 instance. The DB and the application were on the same server. Provisioning an application server was all done with a single bash script. One time we had to re-provision this server and we had take everything down for like 3 hours. We did at 3am so none of our customers got disrupted.
Every single ounce of engineering time was devoted to product development. We ended up getting bought and everyone (all 10 of us; founders were generous and awesome) made a crap ton of money. I can tell you with absolute certainty that every major feature we shipped contributed to that acquisition. Had we dedicated any effort at all to proper infra it just wouldn't have happened. The cost of our crappy duck taped infra was that one night we needed to reprovision everything and maybe another 20 total hours of wasted time across all engineers in the company.
That's why the first line in the blog post specifically cites early stage startups.
Exactly this. Polyrepos, microservices, and their ilk all do the same thing. Bring forward problems you may one day need to solve and make solving them necessary for progress.
Even our background workers ran in the web app. Oh yea, that's right. I can feel the cringing just writing that! So like if we had an expensive background operation, our web app would start up a thread and do the work. Some might be wondering: "What if you needed to restart it? Like, I dunno... when deploying code?". Answer: we didn't, lol. Our deploy script would hold all new jobs and wait until existing ones finished before restarting. Some background process were extremely time consuming (eg: 10+ hours) and those we wrote in a way that they could just resume if killed.
It's actually amazing how productive you can be with a single monolithic web app on a beefy cloud instance.
I was also surprised how quickly we were able to scale out to a proper multi-service architecture after being bought. It only took us 6 months. The big company that bought us plans ahead years. It was no problem at all.
Now don't get me wrong. I'm not advocating for sloppy engineering here. Our team was some of the smartest people I've ever worked with. It wasn't uncommon for me (a medium-experienced engineer) to bring a proposal to our technical founder that would require like 1 month of work. That freakin' genius would reframe the problem and architecture such that we'd spend 2 days, still achieve the same goal, and take on minimal technical debt. I learned A LOT.
I know I'm rambling, but honestly I don't know if that kind of hustle is even a thing anymore. Are tech startups still scrappy? Are they focusing on the core problems/solutions they're trying to prove out instead of layers of architecture? Do such problems even exist anymore? Back in the mid-2000s it felt like that's how everyone did it.
Conway's law is more important: "I need to deploy my team's change without waiting for your team's job to finish".
Sounds mind blowing enough to me :)
I've been there, when one node lost communication with another but not the rest of the cluster - and I've spent some time working with Linux network stack (I mean the actual kernel parts), so I had some clues - but nonetheless I've ended up just scrapping the node with all the workloads and starting a new one. Because I haven't found anything obvious and that was easier than continue debugging. But it wasn't exactly great or fun and painless solution either (I don't exactly remember what sort of trouble I had with draining node but there was something enough to make me swear at the monitor).
So I'm not exactly against K8s but I totally advise to realize its internal complexity (no matter if it comes preconfigured by a cloud provider - it is still there) and ask if one's okay to eventually encounter it.
That's why things like OpenShift exist which cost an order or magnitude more than just a managed control plane.
In the cloud, running workloads on k8s is much more expensive than something simpler, like containers on instances in an AutoScaling group on Amazon, for example. You have to pay for the control plane in addition to the worker nodes, and if you only have a few workers, your control plane costs are significant for small infra. It's fairly easy to set up on {A,E,G}KS.
This cost, however, buys you flexibility. Say your software needs to run on-prem as well as in the cloud? k8s can be your abstraction, and the additional cost becomes worthwhile.
If you use it as a dumb container scheduler, without all this fancy stuff like Persistent Volumes or crazy scheduling constraints, it's not too bad. It does introduce extra complexity around traffic ingress, and is missing some basic functionality, like de-scheduling a Pod from a load balancer, and terminating it once no more connections exist. There are also difficulties around giving people access to the cluster in a safe way (for example, any admin can see the contents of secrets), and kubelets' decision to run without swap also leads to some annoying node failures that would be mere slowdowns if swap was present.
Like anything else, k8s is a mixed bag. If it doesn't solve specific problems for you, don't use it, it's pointless.
As written, none of this is strictly true. I suspect I know why you believe what you wrote is true, but I think you just need to spend some more time understanding the documentation and architecture. As an example, support for swap accounting was added last year, and while swap has not been supported, it's been possible to use swap perhaps forever.
The managed offerings especially don’t let you stand still: they kick off old versions of k8s as they are deprecated.
Here’s my hierarchy of things the average engineer thinks they know well but actually fundamentally never understood and will occasionally screw up badly with escalating consequences as you go down the list:
1. Command line 2. Git 3. eMacs/vim? 4. Docker 5. Kubernetes
I myself personally acknowledge my lack of expertise and try my best to mitigate:
1. This I just try to learn well 2. I only use GitHub desktop (I like it’s limited feature set actually saves you from yourself) 3. Just use py charm? It’s 2022. 4. Docker is necessary but I try to find someone actually decent with it and get them to help me. 5. Just avoid. Unless you have an actual good devops team with 5+ people backing it. Beanstalk or lambda do just fine.