For instance, we had several outages when upgrading the kubernetes version in our clusters. If you have many small cluster it's much easier and more save to apply cluster wide updates, one cluster at a time.
For instance, we had several outages when upgrading the kubernetes version in our clusters. If you have many small cluster it's much easier and more save to apply cluster wide updates, one cluster at a time.
1. Identify one problem you want to fix and ignore everything else.
2. Make a tool to manage the problem while still ignoring everything else.
3. Hype the tool up and shove it in every niche and domain possible.
4. Observer how "everything else" bites you in the ass.
5. Identify the worst problem from #4, use it to start the whole process again.
Microservices will save us! Oh no, they make things complicated. Well, I will solve that with containers! On no, managing containers is a pain. Container orchestration! Oh no, our clusters fail.
Meanwhile, the complexity of our technology goes up and its reliability in practice goes down. Plus, the concepts we operate with are less and less tethered to reality, which makes even talking about certain issues really hard.
This isn't some ageist kids-these-days thing, I mean that the complexity has always existed, and it's always abstracted one level higher over time. Kubernetes doesn't exist because Google and Google-adjacent engineers wanted to foist off "ooh, shiny" on the world, it exists because it solves a problem with containers at scale. Containers solve a problem with resource usage at scale, which are really just an evolution of VMs and a response to their inefficient resource usage and brittleness at scale, and even they were once seen as the new hot fancy overly-complicated thing.
For context, I entered the IT workforce just as virtual machines were being introduced into the average enterprise, and a distressingly high amount of the complaints I hear about containers and k8s were being used against VMs as well.
I'm curious to see where we go from k8s. What the next level of abstraction up will be.
VM/370 came out 49 years ago. So assuming you did a master's at university (I guess math at the time) I'd say you are definite retired by now.
https://en.m.wikipedia.org/wiki/VM_(operating_system)
(yes this is tongue in cheek sort of since you said average but you also said 'enterprise' which my cheek reads as large company :))
... the ancient equivalent of Kubernetes was released in 1973 (JES3)
What I'm playing off of is that the poster probably meant when VMware and such took off in the enterprise. As one would see in my post history I know about what came before and that these things are not new in fact. And in many many cases still in use today which is your second point. What started 49 years ago is still in use on the IBM z series and all of that stuff has awesome backwards compatibility. My dad started his career with 360 assembler (as can also be seen in my post history).
So I guess it's ageism in the direction of _young_ people. I hope it's them down voting. Otherwise it's misjudging peoples sense of humour ;)
I don't do anything at scale. I don't have to scale. My core competency is as a library or tool builder, and often my primary deliverable is a tool or library.
In the past year I got a new JR engineer on my team who was all hot and bothered with docker and k8s and he spent a month changing out CI process to be docker based. It went from a 20 line shell script to a pile of garbage.
I disagreed with the decision at the time, and while a subject matter expert I'm not the team lead so I couldn't say no stop that's bad.
While I'm sure k8s solves problems for Google I'm not google. My company isn't google, and my team isn't solving that class of problems so docker is useless crap for me.
The person is no longer on the team. They left their docker mark, and ran off to dockerify some other project leaving a team with no container expertise and a CI pipeline that is hacks around docker bugs.
The existence of right tools for the job implies the existence of wrong tools for the job. The engineer in GP's story used the wrong tool for the job. That is a people problem. Had he used the right tool for the job, the problem would not have existed. GP wrote their CI stuff in shell to solve a tech problem, not a people problem.
I don’t agree with this at all. Reliability in practice is improving at a phenomenal pace. 10 years ago maintenance outages were a normal feature of every service, and unplanned outages were perfectly ordinary occurrences. Consumers today expect a much higher level of availability and reliability, which they receive rather consistently. Today it takes much less resource to produce a much more reliable system then you would have been able to produce in the rather recent past.
It might take less resource today though to achieve the same, I agree with that.
10 years ago I was working for a company that provided a financial OLTP service. We had to invest a huge amount of money to be able to provide a reasonable HA architecture, and to be able to meet 4 hour DR SLAs, and we still had weekly maintenance outages. The amount of effort required to accomplish those service levels today is comparatively trivial, and you could reasonably expect even a low-budget one person operation to be able to exceed them.
You’d expect a service outage to be a significant public controversy today for a lot of companies. It’s never been a good thing, but we’ve come a long way from it being a completely routine event for most services. Especially given the explosion in online services.
Having said that, in reality that scale applies only to a small minority of products. I think an additional part of the real problem is the ever-present Cargo-Cult Oriented Programming model we've had for a while. I am sure container orchestration is important and maybe even required at places at Google scale. But I hear far too much about its use; it seems that nearly everyone is using it, and I really don't understand why. There seems to be a bit of a keeping-up-with-the-Jones effect, like people would be embarrassed to say that their startup relies on platforms that aren't the latest and greatest. Picking technologies because they sound cool on a resume or at a tech conference, instead of their use being appropriate for the circumstances, seems like a common issue.
The industry strongly incentivises individual engineers to make decisions that will ensure the latest buzzwords appear on their CVs - far more strongly than it incentivises making sound decisions for their current organisation.
There's also very much a tendency to make things complicated in many workplaces as a form of job security and gatekeeping. But it doesn't last forever. Today's hot-shit is tomorrow's ball-and-chain.
10+ years from now, many k8s monsters will still be running and kept alive by stressed-out teams in India. Much like the "Oracle Enterprise Business Suite" giant-shit-ball from the late 90's is still working in many corporations doing absolute critical stuff with tentacles in every part of the org. Changes in these systems are nearly impossible because of the cost and nightmarish complexity, and _everyone_ is _forced_ to use it.
It’s not only that it’s also boredom. Cranking out another cookie-cutter app can be much more fun if you do it in a novel way. Basically all the incentives are aligned towards exotic over-engineering rather than boring, safe but actually perfectly good choices. The danger will come if any savvy manager figures out what this costs them. So we’re probably all safe but you never know.
https://github.com/kubernetes-sigs/multi-tenancy/tree/master...
- C. A. R. Hoare
People would be surprised at how simple the internal infra is, given the fleet size, compared to stuff like k8s.
(I'm talking about the infra that runs on bare metal, not AWS)
Or the infra underlying AWS itself?
And AWS, with its design by accretion, makes that an impossibility. Kubernetes by comparison is a paragon of clarity.
As an AWS customer, I find AWS a huge pile of overcomplicated offerings that is dense with its own jargon and way of doing things. IAM is a trash fire. Everything built on IAM like IRSA is also a trash fire. Why are there managed worker nodes, spot managed worker nodes, Fargate worker nodes, self-managed worker nodes, and spot self-managed worker nodes? Why are there circular dependencies for every piece of infra I stand up making it almost impossible to delete any resources? Why can't I click on a service and see how much it costs me in one or two clicks?
This is a completely insane way of building software to me and I have to eat it.
Amazon has "Invent and Simplify" as one of the leadership principles. However, each team has its own hiring bar and culture, so some will take it seriously and some will not.
> I find AWS a huge pile of overcomplicated offerings
I fully agree. I'm speaking exclusively about the internal infrastructure, not about AWS.
Unfortunately, all my clusters are either confined locally anyway, or are deliberately supposed to be separated in all ways :(
It lets you declaratively define helm deployments, which in turn orchestrate k8s infrastructure.
I dont like where this is going
If you want to have a strong consistent multi-region database (or other form of storage), then you have to deal with slow writes, because your transaction is only allowed to terminate successfully when all masters have written the new data and are in a synchronous state.
I very much prefer slow strong consistent writes instead of fast eventual consistent writes.
Disclaimer- Google employee who is heavily using Spanner
But, whats wrong with postgres here?
For your main production load at work, paying management fees for many clusters is not that relevant costs. So not an issue with my clients' projects.
But for my many hobby-projects, dev and few self-employed business workloads balancing them all on as few clusters as possible is a major pain.
Before that I/you scaled it out to separation of concerns, maybe only grouping a few workloads. And then quite happily scrapping whole clusters all the time.
Now any lifecycle issues with them affects lots of independent applications. They went from cattle to pets over night.
[1] https://www.reddit.com/r/kubernetes/comments/fdgblk/google_g...
What do you do if your Raft data is in zookeeper and you want to move it to etcd or consul? We need 'migrations' for our non-db data too.