How Kubernetes and Kafka Will Get You Fired
medium.com
medium.com
The "vendor agnostic" approach was the right call to make at that point in time. Sure, every business is unique and some cases fit better for a hands-off (lets pay for the convenience of AWS taking care of managed offerings) approach. But, it is a fallacy to think there is no cost to pay for that decision.
The cost of operating your vendor agnostic infrastructure is replaced by your team now needing to learn the intricacies of AWS such as IAM, AWS's way of networking, backups etc. Those operational needs dont just go away, they just become "easier" and more defined as AWS's way of doing it.
As a consultant, one must know where to draw the line and recommend the appropriate route.
Choosing to be vendor agnostic is purposely choosing to need to implement every single part of a stack yourself, rather than the (usually considerably better) AWS offerings, just in case you may want to switch away in the future. It's a massive waste of money and time for nearly every business.
Unless you're in the business of providing an infrastructural-level service (like heroku), you'll ship faster and cheaper by using a single cloud and going all in.
However, once you start growing and have the scale where you have dedicated teams of engineers operating your infrastructure, at some point it makes (economical and operational) sense to own your infra. For example, Instead of 10 different teams each having their own MSK setups to get a kafka experience spinning up so many instances for that occasional need, you could have a centralised kafka cluster operated by a central entity within the team. Instead of every team spinning up their own EKS clusters, and all the cruft associated with making it production ready, the company could set one way of running production services (and have a plan to be able to replicate that outside of AWS should there be a future need to).
> K8s is effectively just a layer over EC2 Yes, but one that comes highly abstract and lures you into tooling such as ingress, secret mgmt, image registry and all kubernetes artifacts that you need to care about in addition to your service itself.
> what's the point of using AWS if all you're using is EC2? Not needing to operate a data center, networking, worrying about needing to visit the server to change the failing disk, and largely - being able to scale up/down on demand without capital expenses.
If I read what I wrote, perhaps it sounds extreme and my grey beard agrees to the ideology more than most HN audience would.
If you have multiple teams spinning up MSK, then you have a different problem and running your own infra isn't going to make this easier. Same thing with EKS clusters. This is a problem with not having a team that manages your infrastructure platform.
Product engineers should own the software they write, should own deploying them (but not the infra that deploys them), should own the oncall (but not the observability infra), and should have access to relevant platforms they need.
Platform engineers should build the platform that platform engineers use. Usually this does not include the CI infrastructure, but in some cases it should.
SRE often owns the observability platform (but in some cases there's an observability team for this).
This is more the norm in large size companies, and usually those teams are split down into multiple other teams that own different parts of the infra that platform engineers depend on for their services. This is usually the case whether or not the company uses a cloud or runs their own infra.
When you run your own infra, you have to have expertise in all the things that you were previously using the cloud for. IMO running k8s yourself is asking for trouble, and EKS is saving you a lot of time and money. Running your own kafka is basically masochistic. Running your own database infrastructure is also extremely difficult, and you'll almost certainly have worse uptime than if you used Aurora (or DynamoDB).
Though you can save money running your own infra, you're going to have to hire people to run it, and you're going to have to plan further out, and will be less agile.
It's very high quality, gives great audit, you can have the same control plane in a ton of regions... EC2 is fantastic.
I don’t have to pay someone else to handle this, but if did, I would get rid of k8s in a heartbeat. I’ve seen a devops team of only a few people manage tens of thousands of traditional servers, but I doubt such a small team could handle a k8s cluster of the same size.
I’m considering moving back to traditional architecture for my blog and other projects. K8s has been fun, but there’s too much magic everywhere.
I’ve heard it’s supposed to solve the problem of programs running differently on different machines. That’s a problem I’ve never encountered in my 12 years of experience.
But the types of issues you describe are very real and very time consuming.
It makes sense if you use docker. Docker containers need somewhere to live. If you want two copies of your service alive at all times, K8s is the thing which will listen for crashes, and restart them, etc.
I already mentioned K8s as an automatic container runner/restarter. But if you run two copies of a service, you need a load balancer to route traffic to them. You can program your own (more work), or download & run someone else's (less work). Or you can see what K8s provides [0] and do even less work than that.
If your services talk to one another, they could talk by hard-coded IP (maintenance nightmare), or by hostname. If they talk by hostname, then they need DNS to resolve those host names. Again, you can roll DNS yourself, or you can see what K8s gives you [1].
And on and on. Firewalls, https, permissions, password/secrets management.
There's one more thing to say about K8s which is that it has become a bit of a defacto standard. So you don't need to relearn a completely new way of doing this stuff if you decide to switch jobs / cloud providers.
[0] https://kubernetes.io/docs/concepts/services-networking/ingr... [1] https://kubernetes.io/docs/concepts/services-networking/dns-...
You seem to also be implying that by running a single bare metal server you have eliminated any chance of downtime which isn’t true
For example if your process crashes on bare metal, you go down, unless you have some kind of supervisor that watched and restarts the process, if youre not using kubernetes as a supervisor then you need to set one up using some other tool. At the end of the day you can’t eliminate all tooling/downtime.
If a single server fails, you may be offline but there are well-tread paths to come back online. Your material cost is the cost of that single server. If k8s goes down, oh boy. Not only is it very complex, requiring knowledge of how it works to diagnose and recover from, but there can be zero documentation on how to recover. You are now also paying for a cloud of bricks.
Or just apply their official helm chart… and you’re pretty much done. You’ll also get better efficiency because the various environments will all get bin-packed together.
Is it perfect? No, but it’s better than doing to yourself!
If you do not need this features k8s is not the thing to use, unless you have the skill set anyways.
Things get messy when state and persistence is involved, I’d prefer to habe my backend DB not on k8s and link the services against it.
To each his own, I guess.
I can definitely see the appeal of tinkering with 'advance tech' for personal hobby tho. Because now I am pretty sure you know more about K8 than me :)
This has been my experience with a lot of the "we need to be cloud native! containers!" mantra in the enterprise. Some exec gets it in their head it's a good idea (and probably gets non-trivial "referral agent fees") this is a must do, and all of the young, hip developer types are happy to cheerlead it.
Two years later OpEx is exploding, most of the processes haven't yet been converted to be in the cloud, and the environment isn't noticeably better or different. It sucks, just sucks in a new and more expensive way that gives you less control of your data.
Seen this at 3 x F500 orgs and with multiple cloud providers, including the big 3 + one of the well known second tiers.
> I spend at least one night a week babysitting the cluster
...
> K8s has been fun
This is why everything sucks now.
Works well with terraform and extending with nomad clients it's a breeze.
I set my personal bare metals with nomad infra and never looked back.
Honestly I didn't got into the position where this to become critical and always managed to stay ahead.
So far(2yrs) so good. I am mostly one man show and I found k8s a bit too much, there is always something.
Maybe I was doing smth wrong or didn't knew how to plan better, idk :)
Edit//typos
Low key I hate touching Google-created projects. On paper technically sound but in practice a guaranteed usability disaster.
Similarly, any cloud provider worth it’s salt provides a managed k8 cluster as a commodity these days.
We got a lot of critical infra running on them and then slowly there was tech-debt that would start accumulating. Clusters have to get updated , older DNS versions in k8s are slow, networking (Older Weave versions was bursting through the seams when the traffic exploded with many applications onboarded). SRE teams get overwhelmed, constant requests for adding PVC (Kafka & C* was on k8s) took a toll. Sanity prevailed in the end, there was decision to move to hosted PaaS infra, though I no longer work there, I just reminisced what we were going through.
Though a "cloud-independent" solution will save pennies, it will definitely drown dollars in personnel costs and the uptime/SLA
History repeats itself, because we don't learn from our mistakes (us or others)
If your team is a bunch of software devs, you are doing the right thing… because k8s requires a bunch of knowledge you likely don’t have if you don’t have a Linux guru on the team. Even if you are using an expensive managed solution, things will go wrong and that knowledge is needed to prevent downtime.
New shop: lambda and eventbridge = life is good.
Very interesting. I am struggling to articulate this mentally. Would you please expand on this comment?
Why didn’t the customer use EKS?
Everytime I am trying to point out systems are more and more complicated and do a call for simplicity I am pushed away.
I heard you are not taken seriously if you don't use a well established cloud provider or similar things.
Truth be told, there are not many projects you do or will work on which need this kind of things.
We though cloud will help us with a lot of things but to what cost and by cost I mean stress, data protection, money etc.
We run Rancher across a couple of bare metal clusters and it's been mostly an amazing experience (ca. 3 years). The only issues we had were with Rancher specific bugs, but those have been resolved and for the most part our infra is pretty autonomous. We do all HA at the application layer, so local NVME as opposed to network storage. This means Patroni, Redis Sentinel/Cluster, etc. But it broadly just works. Maybe we're not big enough to bump into issues, but I couldn't imagine migrating to the labyrinth of vendor lockin masquerading as cloud services.
What am I missing? Why do we have such a wildly different experience to others?
It feels like there's something deeper hiding here, more along the lines of "our developers really don't / can't care about how the software is operating in production."