How we migrated onto K8s in less than 12 months
figma.com
figma.com
I was running on dedicated servers before. My stack is quite complicated and deploys were a nightmare. In the end the dread of deploying was slowing down the little company.
Learning and moving to k8s took me a month. I run around 25 different services ( front ends, product admins, logistics dashboards, delivery routes optimizers, orsm, ERP, recommendation engine, search, etc.... ).
It forced me to clean my act and structure things in a repeatable way. Having all your cluster config in one place allows you to exactly know the state of every service, which version is running.
It allowed me to do rolling deploys with no downtime.
Yes it's complex. As programmers we are used to complex. An Nginx config file is complex as well.
But the more you dive into it the more you understand the architecture if k8s and how it makes sense. It forces you to respect the twelve factors to the letter.
And yes, HA is more than nice, especially when your income is directly linked to the availability and stability of your stack.
And it's not that expensive. I lay around 400 usd a month in hosting.
I'm a K8S believer, but it _is_ complicated. It solves hard problems. If you're multi-cloud, it's a no brainer. If you're doing complex infra that you want a 1:1 mapping of locally, it works great.
But if you're less than 100 developers and are deploying containers to just AWS, I think you'd be insane to use EKS over ECS + Fargate in 2024.
ECS+Fargate however does not. If you are a small team managing the entire stack, you need to factor this into accounts. For example EKS forces you to upgrade the cluster to keep in the main kubernetes release cycle, albeit you can delay it somewhat.
I personally run k8s at home and another two at work and I recommend our teams to use ECS+Fargate+ALB if it is enough for them.
Just use a managed K8s solution that deals with this? AKS, EKS and GKE all do this for you.
This sort of "just" thinking is a great way for teams to drown in ops toil.
https://helm.sh/docs/topics/version_skew/
Istio: https://istio.io/latest/docs/releases/supported-releases/#su...
Literally every kubernetes manifest that hits the server uses a k8s api:
apiVersion: apps/v1(genuine tone, not rhetorical)
But, to reiterate, everything uses APIs. The *betavX APIs are of course likely to be deprecated and replaced with stable APIs after a few versions.
So if you are going to compare with a managed solution, compare with something equivalent. Take a bare managed cluster and add a single Deployment to it, it will be no more complex than ECS, while giving you much better developer ergonomics.
In my experience, most k8s deployments are just "dumb" docker images, they're not very "k8s native".
Your use case may be more complex, which is why you have had more difficulty keeping things up-to-date.
It's a complicated mess compared to something like a Nomad jobspec. That's one of the reasons we decided on Nomad while I was at Cloudflare.
It's been a while since I spun up a k8s instance on AWS, Azure, or the like, but when I did I was astounded at how many implementation decisions and toil I had to do myself. Hosted k8s should be plug-and-play unless you have a very specialized use-case.
Last I checked, managed k8s clusters weren't much more expensive than the compute they ran on.
It's a complicated mess compared to something like a Nomad jobspec. That's one of the reasons we decided on Nomad while I was at Cloudflare.
How do you handle the lack of multi tenancy in Kubernetes?
Happy to hear counterexamples, though — maybe the “indent 4” insanity and multi-level string templating in Helm is gone nowadays?
I'm not a fan of Helm either though, templat-ed yaml sucks, you still have the "indent 4" insanity too. Kustomize is nice when things are simple, but once your app is complex Kustomize is worse than Helm IMO. Try to deploy an app that has a ING, with a TLS cert and external-DNS with Kustomize for multiple environments; you have to patch the resources 3 times instead of just have 1 variable you and use in 3 places.
Helm is popular, Terraform is popular so they both are talked a lot, but IMO there is a tool that is yet to become popular that will replace both of these tools.
For my setup anything that needs to be variable or secret gets specified in a custom json/yaml file which is read by a plugin which in turn outputs the rendered manifest if I can't write it as a "patch". That way the CI/CD runner can access things like the resolved secrets for production without being accessible by developers without elevated access. It requires some digging but there are even annotations that can be used to control things like if Kustomize should add a hash suffix or not to ConfigMap or Secret manifests you generate with plugins.
- it is not really any more lines - doesn’t break if dev upgrades to a different version of the resource (has happened before) - allows you to experiment with dev with other setups (eg additional ingresses, different paths etc) instead of changing a base config which will impact other envs
TLDR patch things that are more or less the same in each env; create complete resources for things that change more.
There is a bit of duplication but it is a lot more simple (see ‘simple made easy’ - rich hockey) than tracing through patches/templates.
However it is loaded with so many footguns that I spend my time redoing and debugging others engineers work.
I’m hoping this new tool called « timoni » picks up steam. It fixes pretty every qualm I have with helm.
So if like me you’re looking for a better solution, go check timoni.
I'm giving it a try and I don't despise it yet, but it feels gross - application configs are typically far more mutable and dynamic than cloud infrastructure configs, and IME, terraform does not likey super dynamic configs.
Best practice as I can currently see it is to have Terraform set up what you need for continuous delivery (e.g. ArgoCD) as part of the infrastructure, then use the CD tool to handle day-to-day deployments. Most CD tooling then asks you to package your deployment in something like Helm.
In our case we have a single application/service base helm chart that provides sane defaults and all our deployments extend from. The amount of helm values config required by the consumers is minimal, and there has been very little occasion for a consumer to include their own templates - the base chart exposes enough knobs to avoid this.
When it comes to third-party charts, many we've been able to deploy as is (sometimes with some PRs upstream to add extra functionality), and occasionally we've needed to wrap/fork them. We've deployed far more third-party charts as-is than not though.
One thing probably worth mentioning w.r.t to maintaining our custom charts is the use of helm unittest (https://github.com/helm-unittest/helm-unittest) - it's been a big help to avoid regressions.
We do manage a few kubernetes resources through terraform, including Argocd (via the helm provider which is rather slow when you have a lot of CRDs), but generally we've found helm chart deployed through Argocd to be much more manageable and productive.
At our company we have all deployments wrapped into a flat helm chart with as little variables as possible. (I always have to fight for that because devs like to abstract helm 100 levels and end up with nothing)
Honestly, as it stands, I think we'd be seen as pretty useless craftsmen in any other field due to an unhealthy obsession of our tooling and meta-work - consistently throwing any kind of sensible resource usage out of the window in favor of just getting to work with certain tooling. It's some kind of a "Temporarily embarrassed FAANG engineer" situation.
In the software/tech industry it's common place to just accept that your app can't be down for any amount of time no matter what. No one checked to see how much more it would cost (engineering time & infra costs) to deploy the app so it would be HA, so no one checked to see if it would be worth it.
I blame this logic on the low interest rates for a decade. I could be wrong.
But a lot of companies are building distributed systems purely because they want this ultra-low downtime. Distributed systems are HARD. You get an entire set of problems you don't get otherwise, and the complexity explodes.
Often, in my opinion, this is not justified. Saving a few minutes of downtime in exchange for making your application orders of magnitude more complex is just not worth it.
Distributed systems solve distributed problems. They're overkill if you just want better uptime or crisis recovery. You can do that with a monolith and a database and get 99.99% of the way there. That's good enough.
But for some deeper details, I'd suggest checking out the comments in this reddit thread[0] (as well as some of the linked articles therein).
E.g. From a comment by /u/Golden_Age_Fallacy: A great use of Nomad is on reduce the burden of on-boarding a team(s) of developers who are unfamiliar with cloud native deployments / systems(even containers!).
Nomad jobspecs are very simple and straight forward, as compared to the complexity and pure option overload you get in k8s and helm.
From /u/neutralized: It's much easier to use than k8s. Easy to setup, easy to manage, much more shallow learning curve. Nothing super fancy. Just works. I migrated a startup I was at off of a self-managed k8s setup to Nomad a few years ago and they've never looked back.
From /u/esity: My team is currently building out a fully automated nomad cluster service offering internally(fortune 10)
It's super awesome. Easy. Little headache. Integrates with consul and vault. We are literally planning to replace thousands of vms for K8s with nomad. Containers are faster, more resilient and writing hcl is actually fun once you learn it
Now, there is a rather more lengthy comment, by /u/thomasbuchinger, that goes through the pros and cons he experienced in trying Nomad out and his conclusion is that, while he wouldn't discourage anyone from using it, "k3s and a few well-known simple projects give you 80% of Nomands [sic] features. Are as easy to operate, afford you more options in the future and have a ton of documentation/tutorials...available."
There are more comments in the thread and again links to a bunch of blogposts/articles/etc., including one from fly.io that seemed pretty detailed, discussing the Googly origins of both k8s and Nomad (fly.io used Nomad but found that it wasn't the best fit for them, which is also discussed in their post -- actually, I'm going to put the link to their post below[1], since I think it is worthwhile).
Hope all this helps.
[0] https://old.reddit.com/r/devops/comments/11nsxo3/opinions_on...
[1] https://fly.io/blog/carving-the-scheduler-out-of-our-orchest...
While nomad can be interesting again it's a single "smallish" vendor pushing an "open" (see debacle with Teraform) source project.
FAANG engineers made the same mistake, too, even though the analogy implies comparative competency or value.
In some ways, it's nice to see companies move to use mostly open source infrastructure, a lot of it coming from CNCF (https://landscape.cncf.io), ASF and other organizations out there (on top of the random things on github).
There are kata-containers I think, they might solve my angst and make me enjoy k8s
Overall... There's just nothing cool in kubernetes to me. Containers, load balancers, megabytes of yaml -- I've seen it all. Nothing feels interesting enough to try
If you have never dealt with, I have to run these 50 containers plus Nginx/CertBot while figuring out which node is best to run it, yea, I can see you not being thrilled about Kubernetes. For the rest of us though, Kubernetes helps out with that easily.
if there's a kernel vulnerability in something simple (like dirtycow, which was if I remember correctly about pipes) then the attacker will take over your entire 128 core machine and all the hundreds applications there
I don't want to set up or maintain my own ELK stack, or Prometheus. Or wrestle with CNI plugins. Or Kafka. Or high availability Postgres. Or Argo. Or Helm. Or control plane upgrades. I can get up and running with the AWS equivalent almost immediately, with almost no maintenance, and usually with linear costs starting near zero. I can solve business problems so, so much faster and more efficiently. It's the difference between me being able to blow away expectations and my whole team being quarters behind.
That said, when there is a genuine multi-cloud or on-prem requirement, I wouldn't want to do it with anything other than k8s. And it's probably not as bad if you do actually work at a company big enough to have a lot of skilled engineers that understand k8s--that just hasn't been the case anywhere I've worked.
Because as annoying as getting the prometheus + loki + tempo + promtail stack going on k8s is —- I don’t really believe that writing it from scratch is easier.
* Log aggregation happens out of the box with CloudWatch Logs and CloudWatch Log Insights. It's configurable if you want different behavior
* On ECS, you configure a "service" which describes how many instances of a "task" you want to keep running at a given time. It's the abstraction that handles spinning up new tasks when one fails
* ECS supports ready checks, and (as noted above) integrates with ALB so that requests don't get sent to containers until they pass a readiness check
* Machine maintenance schedules are non-existent if you use ECS / Fargate, or at least they're abstracted from you. As long as your application is built such that it can spin up a new task to replace your old one, it's something that will happen automatically when AWS decommissions the hardware it's running on. If you're using ECS without Fargate, it's as simple as changing the autoscaling group to use a newer AMI. By default, this won't replace all of the old instances, but will use the new AMI when spinning up new instances
But again, though: the biggest selling point is the lack of maintenance / babysitting. If you set up your stack using ECS / Fargate and an ALB five years ago, it's still working, and you've probably done almost nothing to keep it that way.
You might be able to do the same with Kubernetes, but your control plane will be out of date, your OSes will have many missed security updates. Might even need a major version update to the next LTS. Prometheus, Loki, Tempo, Promtail will be behind. Your helm charts will be revisions behind. Newer ones might depend on newer apiVersions that your control plane won't support until you update it. And don't forget to update your CNI plugin across your cluster, too.
It's at least one full time job just keeping all that stuff working and up-to-date. And it takes a lot more know-how than just ECS and ALB.
Most of my managed Kubernetes experience is through Amazon's EKS, and the pain I remember included frustration from the supported Kubernetes versions being behind the upstream versions, lack of visibility for troubleshooting control nodes, and having to explain / understand delays in NIC and EBS appropriation / attachments for pods. Also the ALB ingress controller was something I needed to install and maintain independently (though that may be different now).
Though that was also without us going neck-deep into being vendor agnostic. Using EKS just for the Kubernetes abstractions without trying hard to be vendor agnostic is valid--it's just not what I was comparing above because it was usually that specific business requirement that steered us toward Kubernetes in the first place.
If you ARE using EKS with the intention of keeping as much as possible vendor agnostic, that's also valid, but then now you're including a lot of the stuff I complained about in my other comment: your own metrics stack, your own logging stack, your own alarm stack, your own CNI configuration, etc.
- ALBs -- yeah this is correct. However ALBs have much longer startup/health check times than Envoy/Traefik
- Cloudwatch - this is true, however the "configurable" behavior makes cloudwatch trash out of the box. you get i.e. exceptions split across multiple log entries with the default configure
- ECS tasks - yep, but the failure behavior of tasks is horrible because there're no notifications out of the box (you can configure it)
- Fargate does allow you to avoid maintenance, however it has some very hairy edges like i.e. you can't use any container that expects to know its own ip address on a private vpc without writing a custom script. Networking in general is pretty arcane on Fargate and you're going to have to manually write and maintain the breakages from all this
> You might be able to do the same with Kubernetes, but your control plane will be out of date, your OSes will have many missed security updates. Might even need a major version update to the next LTS. Prometheus, Loki, Tempo, Promtail will be behind. Your helm charts will be revisions behind. Newer ones might depend on newer apiVersions that your control plane won't support until you update it. And don't forget to update your CNI plugin across your cluster, too.
I think maybe you haven't used K8S in years. Karpenter, EKS, + a GitOps (Flux or Argo) makes you get the same machine maintenance feeling as ECS but on K8S without any of the annoyances of dealing with ECS. All your app versions can be pinned or set to follow latest as you prefer. You get rolling updates each time you switch machines (same as ECS, and if you really want to you can run on top of Fargate).
By contrast, if your ECS/Fargate instance fails you haven't mentioned any notifications in your list -- so if you forgot to configure and test that correctly, your ECS could legitimately be stuck on a version of your app code that is 3 years old and you might not know if you haven't inspected the correct part of amazon's arcane interface.
By the way, you're paying per use for all of this.
At the end of the day, I think modern Kubernetes is strictly simpler, cheaper, and better than ECS/Fargate out of the box and has the benefit of not needing to rely on 20 other AWS specific services that each have their own unique ways of failing and running a bill up if you forget to do "that one simple thing everyone who uses this niche service should know".
Sure there are a lot of great feature with k8s which ECS cannot do, but when ECS does satisfy the requirements, it will require less maintenance, no matter what kind of k8s you compare it against to.
Obviously that part's not different from Kubernetes, but here's the part that is: maintenance and upgrades are either completely out of my scope or absolutely minimal. On ECS, it might involve switching to a more recently built AMI every six months or so. AWS is famously good about not making backward incompatible changes to their APIs, so for the most part, things just keep working.
And don't forget you'll need a lot of those AWS skills to run Kubernetes on AWS, too. If you're lucky, you'll get simple use cases working without them. But once PVCs aren't getting mounted, or pods are stuck waiting because you ran out of ENI slots on the box, or requests are timing out somewhere between your ALB and your pods, you're going to be digging into the layer between AWS and Kubernetes to troubleshoot those things.
I run Kubernetes at home for my home lab, and it's not zero maintenance. It takes care and feeding, troubleshooting, and resolution to keep things working over the long term. And that's for my incredibly simple use cases (single node clusters with no shared virtualized network, no virtualized storage, no centralized logs or metrics). I've been in charge of much more involved ones at work and the complexity ceiling is almost unbounded. Running a distributed, scalable container orchestration platform is a lot more involved than piggy backing on ECS (or Lambda).
It's the ecosystem around it which turns things needlessly complex.
Just because you have kubernetes, you don't necessarily need istio, helm, Argo cd, cilium, and whatever half baked stuff is pushed by CNCF yesterday.
For example take a look at helm. Its templating is atrocious, and if I am still correct, it doesn't have a way to order resources properly except hooks. Sometimes resource A (deployment) depends on resource B (some CRD).
The culture around kubernetes dictates you bring in everything pushed by CNCF. And most of these stuff are half baked MVPs.
---
The word devops has created expectations that back end developer should be fighting kubernetes if something goes wrong.
---
Containerization is done poorly by many orgs, no care about security and image size. That's a rant for another day. I suspect this isn't a big reason for kubernetes hate here.
For instance: they want to go to k8s because they want to use etcd/helm, which they can't on ECS? Why do you want to use etcd/helm? Is it really this important? Is there really no other way to achieve the goals of the company than exactly like that?
When a decision is founded on a desire of the user, its easy to validate that downstream decisions make sense. When a decision is founded on a technological desire, downstream decisions may make sense in the context of the technical desire, but do they make sense in the context of the user, still?
Either I don't understand organizations of this scale, or it is fundamentally difficult for organizations of this scale to identify and reason about valuable work.
Isn't it slightly depressing that this explanation is fairly (the most?) plausible?
This was just a Microsoft licensing negotiation tactic. Before he was CEO, Ballmer flew here to negotiate one of the contracts. The discounts were epic.
Many are so-analyzed, but usually in ways that anyone who paid attention in high school science or stats classes can tell are so flawed that they’re meaningless.
We can’t even measure manager efficacy to any useful degree, in nearly all cases. We can come up with numbers, but they don’t mean anything. Good luck with anything more complex.
Very small organizations can probably manage to isolate enough variables to know how good or bad some move was in hindsight, if they try and are competent at it (… if). Sometimes an effect is so huge for a large org that it overwhelms confounders and you can be pretty confident that it was at least good or bad, even if the degree is fuzzy. Usually, no.
Big organizations are largely flying blind. This has only gotten worse with the shift from people-who-know-the-work-as-leadership to professional-managers-as-leadership.
Being vendor-locked into ECS means you must pay whatever ECS wants... using k8s means you can feasibly pick up and move if you are forced.
Even if it doesn't save money today it might save a tremendous amount in the future and/or provide a much stronger position to negotiate from.
Just like picking EKS you have to be aware of the pros and cons of picking the cloud provider tool or not. Luckily the CNCF is doing a lot for reducing vender lock in and I think it will only continue.
1. The time it will take to move to another cloud is proportional to the complexity of your app. For example, if you're a Go shop using managed persistence are you more vendor locked in any meaningful way than k8s? What's the delta here?
2. Do you really think you can haggle with the fuel-producers like you're MAERKS? No, you're more likely just a car driving around for a gas station with increasingly diminishing returns.
It's even worse when your entire platform is vendor-locked.
There is nothing but upside to working towards a vendor-neutral position. It gives you options. Even if you never use those options, they are there.
> Do you really think you can haggle
At the scale of someone like Figma? Yes, they do negotiate rates - and a competent account manager will understand Figma's position and maximize the revenue they can extract. Now, if the account rep doesn't play ball, Figma can actually move their stuff somewhere else. There's literally nothing but upside.
I swear, it feels like some people are just allergic to anything k8s and actively seek out ways to hate on it.
[1] https://auth0.com/blog/upcoming-pricing-changes-for-the-cust...
Most people looking into (and using) k8s that are being told the "you most avoid vendor lock in!" selling point are nowhere near the size where it matters. But I know there's essentially bulk-pricing, as we have it where I work as well. That it's because of picking k8s or not however is an extremely long stretch, and imo mostly rationalization. There's nothing saying that a cloud move without k8s couldn't be done within the same amount of time. Or that even k8s is the main problem, I imagine it isn't since it's usually supposed to be stateless apps.
Where you buy compute from is just as big of a deal as where you buy your other SaaS' from. In all of the cases, if you cannot move even if you had to (ie. it'll take 1 year+ to move), then you are not in a good position.
Addressing your #1 point - if you use a regular database that happens to be offered by a cloud provider (ie. Postgres, MySQL, MongoDB, etc) then you can pick up and move. If you use something proprietary like CosmoDB, then you are stuck or face significant efforts to migrate.
With k8s, moving to another cloud can be as simple as creating an account and updating your configs to point at the new cluster. You can run every service you need inside your cluster if you wanted. You have freedom of choice and mobility.
> Most people looking into (and using) k8s that are being told the "you most avoid vendor lock in!" selling point are nowhere near the size where it matters.
This is just simply wrong, as highlighted by the SaaS example I provided. If you think you are too small so it doesn't matter, and decide to embrace all of the cloud vendor's proprietary services... what happens to you when that cloud provider decides to change their billing model, or dramatically increases price? You are screwed and have no options but cough up more money.
There's more decisions to make and consider regarding choosing a cloud platform and services than just whatever is easiest to use today - for any size of business.
I have found that, in general, people are afraid of using k8s because it isn't trivial to understand for most developers. People often mistakenly believe k8s is only useful when you're "google scale". It solves a lot of problems, including reduced vendor-lock.
K8s makes this really easy. Don't need to worry whether country X has a local Cloud data center of Vendor Y.
Plus it makes hiring so much easier as you only need to understand the abstraction layer.
We don't hire people for ARM64 or x86. We have abstraction layers. Multiple even.
We'd be fooling us not to use them.
I mean the blog post is written by the team deciding the company needs. They explained exactly why they can't easily use etcd on ECS due to technical limitations. They also talked about many other technical limitations that were causing them issues and increasing cost. What else are you expecting?
Aline the VM upgrade, auth, backup, log rotation etc.
With k8s I can give everyone a namespace, policies, volumes, have automatic log aggregation due to demon sets and k8s/cloud native stacks.
Self healing and more.
It's hard to describe how much better it is.
At its core we are a platform teams building tools, often for other platform teams, that are building tools that support the developers at Figma creating the actual product experience. It is often harder to reason about what the right decisions are when you are further removed from the end user, although it also gives you great leverage. If we do our jobs right the multiplier effect of getting this platform right impacts the ability of every other engineer to do their job efficiently and effectively (many indirectly!).
You bring up good examples of why this is hard. It was certainly an alternative to say sorry we can't support etcd and helm and you will need to find other ways to work around this limitation. This was simply two more data points helping push us toward the conclusion that we were running our Compute platform on the wrong base building blocks.
While difficult to reason about, I do think its still very worth trying to do this reasoning well. It's how as a platform team we ensure we are tackling the right work to get to the best platform we can. Thats why we spent so much time making the decision to go ahead with this and part of why I thought it was an interesting topic to write about.
First, when some team says "we want to use helm and etcd for some reason and we haven't been able to figure out how to get that working on our existing platform," start by asking them what their actual goal is. It is obscenely unlikely that helm (of all things) is a fundamental requirement to their work. Installing temporal, for example, doesn't require helm and is actually simple, if it turns out that temporal is the best workflow orchestrator for the job and that none of the probably 590 other options will do.
Second, once you have figured out what the actual goal is, and have a buffet of options available, price them out. Doing some napkin math on how many people were involved and how much work had to go into it, it looks to me that what you have spent to completely rearchitect your stack and operations and retrain everyone -- completely discounting opportunity cost -- is likely not to break even in even my most generous estimate of increased productivity for about five years. More likely, the increased cost of the platform switch, the lack of likely actual velocity accrual, and the opportunity cost make this a net-net bad move except for the resumes of all of those involved.
So am I reading this right that either downstream platform teams or devs wanted to leverage existing helm templates to provision infrastructure and being on ECS locked you out of those and the water eventually boiled over. If so that's a pretty strong statement about the platform effect of k8s.
I understand what you're saying, the thing that worries me though is that the input you get from other technical teams is very hard to verify. Do you intend to measure the development velocity of the teams before and after the platform change takes effect?
In my experience it is extremely hard to measure the real development velocity (in terms of value-add, not arbitrary story points) of a single team, not to mention a group of teams over time, not to mention as a result of a change.
This is not necessarily criticism of Figma, as much as it is criticism of the entire industry maybe.
Do you have an approach for measuring these things?
The canonical way to do that is to ensure that the incoming demand comes with both the ask and also the solid justification. Even at top tier organizations, frequently asks are good ideas, sensible ideas, nice ideas, probably correct ideas -- but none of that is good enough/acceptable enough. The proportion of good/sensible/nice/probably correct ideas that are justifiable is about 5% in my lived experience of 38 years in the industry. The onus is on the asking team to provide that full true and complete justification with sufficiently detailed data and in the manner and form that convinces the platform team's leadership. The bar needs to be high and again, has to provide a clear line of sight to improving the life of the paying customer. The platform team has the authority and agency necessary to defend the customer, operations and their time, and can (and often should) say no. It is not the responsibility of the platform team to try to prove or disprove something that someone wants, and it's not 'pushing back' or 'bureaucracy', it's basic sober purpose-of-the-company fundamentals. Time and money are not unlimited. Nothing is free.
Frequently the process of trying to put together the justification reveals to the asking team that they do not in fact have the justification, and they stop there and a disaster is correctly averted.
Sometimes, the asking team is probably right but doesn't have the data to justify the ask. Things like 'Let's move to K8s because it'll be better' are possibly true but also possibly not. Vibes/hacker news/reddit/etc are beguiling to juniors but do not necessarily delight paying customers. The platform team has a bunch of options if they receive something of that form. "No" is valid, but also so is "Maybe" along with a pilot test to perform A/B testing measurements and to try to get the missing data; or even "Yes, but" with a plan to revert the situation if it turns out to be too expensive or ineffective after an incrementally structured phase 1. A lot depends on the judgement of the management and the available bandwidth, opportunity cost, how one-way-door the decision is, etc.
At the end of the day, though, if you are not making a data-driven decision (or the very closest you can get to one) and doing it off naked/unsupported asks/vibes/resume enhancement/reddit/hn/etc, you're putting your paying customer at risk. At best you'll be accidentally correct. Being accidentally correct is the absolute worst kind of correct, because inevitably there will come a time when your luck runs out and you just killed your team/organization/company because you made a wrong choice, your paying customers got a worse/slower-to-improve/etc experience, and they deserted you for a more soberly run competitor.
All you need is a way to roll out your artifact to production in a roll over or blue green fashion after the preparations such as required database alterations be it data or schema wise.
Easier said than done.
You can start by implementing this yourself and thinking how simple it is. But then you find that you also need to decide how to handle different environments, configuration and secret management, rollbacks, failover, load balancing, HA, scaling, and a million other details. And suddenly you find yourself maintaining a hodgepodge of bespoke infrastructure tooling instead of your core product.
K8s isn't for everyone. But it sure helps when someone else has thought about common infrastructure problems and solved them for you.
And then all you're left with is scaling. Which most business do not need.
Almost everything you've written there is a standard feature of almost any CI toolchain, teamcity, Jenkins, Azure DevOps, etc., etc.
We were doing it before k8s was even written.
Build tools? These are runtime and operational concerns. No build tool will handle these things.
> And then all you're left with is scaling. Which most business do not need.
Eh, sure they do. They might not need to hyperscale, but they could sure benefit from simple scaling, autoscaling at peak hours, and scaling down to cut costs.
Whether they need k8s specifically to accomplish this is another topic, but every business needs to think about scaling in some way.
> Almost everything you've written there is a standard feature of almost any CI toolchain, teamcity, Jenkins, Azure DevOps, etc., etc.
Huh? Please explain how a CI pipeline will handle load balancing, configuration and secret management, and other operational tasks for your services. You may use it for automating commands that do these things, but CI systems are entirely decoupled from core infrastructure.
> We were doing it before k8s was even written.
Sure. And k8s isn't the absolute solution to these problems. But what it does give you is a unified set of interfaces to solve common infra problems. Whatever solutions we had before, and whatever you choose to compose from disparate tools, will not be as unified and polished as what k8s offers. It's up to you to decide the right trade-off, but I find the head-in-the-sand dismissal of it equally as silly as cargo culting it.
https://learn.microsoft.com/en-us/azure/devops/pipelines/pro...
They've always had this stuff. Back in the day it used to be XML transforms. I know azure have added key vault in recent years as well so you can get the values from the key vault instead, but it seems a cynical cash grab to me.
For loading balancing, there's a reason the load balancing servers all have APIs.
Almost no business will save money scaling. Think about it. Most businesses are small and medium business that just don't have the volume of traffic to justify anything more than buying or renting a single/two servers vs the cost of having a scaling solution in place. There are only a small number of businesses (in the scheme of things) that benefit from scaling as you have to pay to do it. And even then the saving is often so miniscule compared to a dev salary you lose money paying people to even set it up and will NEVER recoup that cost.
The cost of even having a single meeting with 4 people about it is more than the saving over a year.
I'm no fan of Helm, but there are surprisingly few good alternatives to etcd (i.e. highly available but consistent datastores, suitable for e.g. the distributed equivalent of a .pid file) - Zookeeper is the only one that comes to mind, and it's a real pain on the ops side of things, requiring ancient JVM versions and being generally flaky even then.
> This meant that something as simple as adding an environment variable required writing and applying Terraform, then running a deploy
This sounds less like a problem with ECS and more like an overcomplication in how they were using terraform + ECS to manage their deployments.
I get the generating templates part for verification prior to live deploys. But this seems... dunno.
To me that seems like the simpler solution.
Unfortunately in my experience, this is true until it isn't. Once it isn't true, it can quickly become a painful blackbox debugging exercise. If your org is big enough to have dedicated AWS support then they can often get help from engineers, but if you aren't then life can get really complicated.
Still not a bad choice for most apps though, especially if it's just a run-of-the-mill HTTP-based app
What a heck am I reading? For who? I am not sure why companies even bother with such migrations. Where is the business value? Where is the gain for the customer? Is this one of those "L'art pour l'art" project that Figma does it just because they can?
There was plenty of services that were latency sensitive or in the HPC realm where it made no sense to force a migration though, and there was no attempt to force them to shoehorn in.
I'm really wondering if I'm aging out of professional software development.
It’s not the only approach so you may well be familiar with others.
Helm, for instance, is a great time saver for installing software. Often software will support nothing but helm. Ease of deployment is a good consideration. Their points on networking are absolutely spot on. The scaling considerations are spot on. Killing/isolating unhealthy containers is completely valid. I could go on a lot more, but I don't see a single point listed as invalid.
One system I can think of off the top of my head is when Amazon moved away from Oracle to fully Amazon/OSS RDBMSs a while ago, but that was multi year I think. If they could have done it in less than a year, they'd definitely be bragging.
A common pattern you'll see though is skipping writing any sort of code and instead using a higher level dsl-ish configuration usually via yaml, using tools like Kyverno.
On the platform integration/hook side (app code doing specialised platform-specific integration stuff, extensions to k8s itself), golang is the lingua franca but bindings for many languages are around and good.
If you want to build your own Kubernetes Custom Resources and Controllers, Go lang works pretty well for that.
Huh? You've been running on AWS for how long and haven't been using auto scaling AT ALL? How was this not priority number one for the company to fix? You're just intentionally burning money at that point!
> While there is some support for auto-scaling on ECS, the Kubernetes ecosystem has robust open source offerings such as Keda for auto-scaling. In addition to simple triggers like CPU utilization, Keda supports scaling on the length of an AWS Simple Queue Service (SQS) queue as well as any custom metrics from Datadog.
ECS autoscaling is easy, and supports these things. Fair play if you just really wanted to use CNCF projects, but this just seems like you didn't really utilize your previous infrastructure very well.
This was actually added beginning of the year. Definitely was on my most wanted list for a while. You could technically use EFS, but that’s a very expensive way to run anything IO intensive.
“ECS doesn’t support helm charts!”
No shit sherlock, that’s a thing literally built on Kubernetes. It’s like a government RFP that can only be fulfilled by a single client.
I dunno, that seems like a very good reason to me.
Unless they are in this space competing against k8s, it’s reasonable for them if they want to use Helm charts, to move where they can.
Also, Helm doesn’t work with ECS, so doesn’t <50 other tools and tech from the CNCF map>.
With k8s you can easily deploy a helm chart that will deploy lots of things that all work together fairly easily.
Y'all really don't think a company like Figma stands to benefit from the flexibility that Kubernetes offers?
At a certain scale, K8s is the simple option.
I think much of the hate on HN comes from the "ruby on rails is all you need" crowd.
Maybe - people seem really gungho about serverless solutions here too
Yes, you could say that. :)
Though you probably could still do it, but that's likely more trouble then it's worth
(The 10k is just an arbitrary number I made up, there is no magic number which makes this approach unviable, it all depends on how the users interact with the platform/how often and where the data is inserted)
Because single hosts will always go down. Just a question of when.
You can view a single k8s also as a single host, which will go down at some point (e.g. a botched upgrade, cloud network partition, or something similar). While much less frequent, also much more difficult to get out of.
Of course, if you have a multi-cloud setup with automatic (and periodically tested!) app migration across clouds, well then... Perhaps that's the answer nowadays.. :)
Kubernetes is a remarkably reliable piece of software. I've administered (large X) number of clusters that often had several years of cluster lifetime, each, everything being upgraded through the relatively frequent Kubernetes release lifecycle. We definitely needed some maintenance windows sometimes, but well, no, Kubernetes didn't unexpectedly crash on us. Maybe I just got lucky, who knows. The closest we ever got was the underlying etcd cluster having heartbeat timeouts due to insufficient hardware, and etcd healed itself when the nodes were reprovisioned.
There's definitely a whole lotta stuff in the Kubernetes ecosystem that isn't nearly as reliable, but that has to be differentiated from Kubernetes itself (and the internal etcd dependency).
> You can view a single k8s also as a single host, which will go down at some point (e.g. a botched upgrade, cloud network partition, or something similar)
The managed Kubernetes services solve the whole "botched upgrade" concern. etcd is designed to tolerate cloud network partitions and recover.
Comparing this to sudden hardware loss on a single-VM app is, quite frankly, insane.
I don’t get it either. It’s not hard at all.
It restarts every single container in the cluster at the same time: https://github.com/kubernetes/kubernetes/issues/122028
We have also found data races in the statefulset controller which only occurs when you have thousands of statefulsets.
Overall, if you stay on the beaten path k8s reliability is good.
You can run rails on a single host using a database on the same server. I've done it and it works just fine as long as you tune things correctly.
Can you elaborate?
- Limiting memory usage and number of connections for mysql
- Tracking maximum memory size of rails application servers so you didn't run out a memory by running too many of them
- Avoid writing unnecessarily memory intensive code (This is pretty easy in ruby if you know what you're doing)
- Avoiding using gems unless they were worth the memory use
- Configuring the frontend webserver to start dropping connections before it ran out of memory (I'm pretty sure that was just a guess)
- Using the frontend webserver to handle traffic whenever possible (mostly redirects)
- Using IP tables to block traffic before hitting the webserver
- Periodically checking memory use and turning off unnecessary services and cronjobs
I had the entire application running on a 512mb VPS with roughly 70mb to spare. It was a little less spare than I wanted but it worked.
Most of this was just rate limiting with extra steps. At the time rails couldn't use threads, so there was a hard limit on the number of concurrent tasks.
When the site went down it was due to rate limiting and not the server locking up. It was possible to ssh in and make firewall adjustments instead of a forced restart.
Personally I think k8s is where it's at now. The innovation and open source contributions are immense.
I'm glad we made the switch. I understand the frustrations of the past, but I think it was much harder to use 4+ years ago. Now, I don't see how anyone could mess it up so hard.
I’ve been struggling to square this sentiment as well. I spend all day in AWS and k8s and k8s is at least an order of magnitude simpler than AWS.
What are all the people who think operating k8s is too complicated operating on? Surely not AWS…
Doing anything in AWS is like pulling teeth.
(Apologies if this is a dumb question) but isn't Figma big enough to want to do any of their stuff on their own hardware yet? Why would they still be paying AWS rates?
Or is it the case that a high-profile blog post about K8S and being provider-agnostic gets you sufficient discount on your AWS bill to still be value-for-money?
Well, that's one hypothesis.
Another is that "Every maturing company with predictable products must be exploring ways to move workloads out of the cloud. AWS took your margin and isn't giving it back." ( https://news.ycombinator.com/item?id=35235775 )
Not managing procurement of hardware, upgrades, etc, and a defined standard operating model with accessible documentation and the ability to hire people with experience, and have to hire less people as you are doing less is enough to build a viable and demonstrable business case.
Scale beyond a certain point is hard without support and delegated responsibility.
I wonder if as time goes on that skill to use hardware is dissappearing. New engineers don't learn it, and the ones that slowly forget. I'm not that sharp on anything I haven't done in years, even if it's in a related domain.
I can imagine the productivity of spinning up elastic cloud resources vs fixed data center resourcing being more important, especially considering how frequently a company like Figma ships new features.
Their ARR in 2022 was around $400M-450M. Say the infra budget at a typical 10% would be $50M. While it is a lot of money, it is not build your hardware money, also not all of it would be compute budget. They also would be spending on other SaaS apps like say Snowflake etc to special workloads like with GPUs, so not all workloads would be in-house ready. I would surprised if their commodity compute/k8s is more than half their overall budgets.
It is lot more likely to slow product growth to focus on this now, especially since they were/are still growing rapidly.
Larger SaaS companies than them in ARR still find using cloud exclusively is more productive/efficient.
They are almost certainly not paying sticker prices. Above a certain size, companies tend to have bespoke prices and SLAs that are negotiated in confidence.
Data centers are wildly expensive to operate if you want proper security, redundancy, reliability, recoverability, bandwidth, scale elasticity, etc.
And when I say security, I'm not just talking about software level security, but literal armed guards are needed at the scale of a company like Figma.
Bandwidth at that scale means literally negotiating to buy up enough direct fiber and verifying the routes that fiber takes between data centers.
At one of the companies I worked at, it was not uncommon to lose data center connectivity because a farmer's tractor cut a major fiber line we relied on.
Scalability might include tracking square footage available for new racks in physical buildings.
As long as your company is profitable, at anything but Facebook like scale, it may not be worth the trouble to try to run your own data center.
Even if the cloud doesn't save money, it saves mental energy and focus.
Renting a few hundred cabinets from Equinix or Digital Realty is going to potentially be hugely cheaper than AWS, but you probably need a team of people to run it. That can be worthwhile if your growth is predictable and especially if your AWS bandwidth bill is expensive.
But then you’re building on bare metal. Gotta deploy your own databases, maybe kubernetes for running workloads, or something like VMware for VMs. And you don’t get any managed cloud services, so that’s another dozen employees you might need.
* Service discovery
* Auto bin packing
* Load Balancing
* Automated rollouts and rollbacks
* Horizonal scaling
* Probably more I forgot about
You also have secret and config management built in. If you use k8s you also have the added benefit of making it easier to move your workloads between clouds and bare metal. As long as you have a k8s cluster you can mostly move your app there.
Problem is most companies I've worked at in the past 10 years needed multiple of the features above, and they decided to roll their own solution with Ansible/Chef, Terraform, ASGs, Packer, custom scripts, custom apps, etc. The solutions have always been worse than what k8s provides, and it's a bespoke tool that you can't hire for.
For what k8s provides, it isn't complex, and it's all documented very well, AND it's extensible so you can build your own apps on top of it.
I think there are more SWE on HN than Infra/Platform/Devops/buzzword engineers. As a result there are a lot of people who don't have a lot of experience managing infra and think that spinning up their docker container on a VM is the same as putting an app in k8s. That's my opinion on why k8s gets so much hate on HN.
But once you start using k8s you probably tend to scope creep and find a lot of shiny things to add to your set up.
# A verbose comment that starts capitalized, followed by a single line of code, cuing you that it was written by a ChatBot.
Some ways to tell if someone is a great developer is hard. You can't tell if someone is a brilliant shipper of features, choosing exactly the right concerns to worry about at the moment, like doing more website authoring and less devops, with a grand plan for how to make everything cohere later; or, if the guy just doesn't know what the fuck he is doing.Kubernetes adoption is one of those, hard ones. It isn't a strong, bright signal like using PEP 8 and having a `pyproject.toml` with dependencies declared. So it may be obvious to you, "People adopt Kubernetes over ad-hoc decoupled solutions like Terraform because it has, in a Darwinian way, found the smallest set of easily surmountable concerns that should apply to most good applications." But most people just see, "Ahh! Why can't I just write the method bodies for Python function signatures someone else wrote for me, just like they did in CS50!!!"
The _minute_ you start running containers in the cloud you need to think of "what happens if it goes down/how do I update it/how does it find the database", and you need an orchestrator of some sort, IMO. A managed service (I prefer ECS personally as it's just stupidly simple) is the way to go here.
I really do think Google Cloud Run/Azure Container Apps (and then in AWS-land ECS-on-fargate) is the right solution _especially_ in that case - you just shove a container on and tell it the resources you need and you're done.
#cloud-config
apt:
sources:
docker.list:
source: deb [arch=amd64] https://download.docker.com/linux/ubuntu $RELEASE stable
keyid: 9DC858229FC7DD38854AE2D88D81803C0EBFCD88
packages:
- docker-ce
- docker-ce-cliInstead you can do those 3 and more in k8s and it would be the same manifests regardless which k8s cluster you deploy to, EKS, AKS, GKE, on prem, etc.
Plus you don't get service discovery across VMs, you don't get a CSI so good luck if your app is stateful. How do you handle secrets, configs? How do you deploy everything, Ansible, Chef? The list goes on and on.
If your app is simple sure, I haven't seen simple app in years.
Also, if you can't convert an ALB into an Azure Load balancer, then you probably have no business doing any sort of software development.
ALB costs get very steep very quickly too, but you're right - start with ALB and then migrate to nginx when costs get too high
If you do this just use GKE autopilot. It’s cheaper and done for you.
For example, even servers (aka instances/vms/vps) with load balancers (aka fabric/mesh/istio/traefik/caddy/nginx/ha proxy/ATS/ALB/ELB/oh just shoot me) in front existed for apps that are LARGER than can fit on a single server (virtually the definition of horizontally scalable). These apps are typically monoliths or perhaps app tiers that have fallen out of style (like the traditional n-tier architecture of app server-cache-database, swap out whatever layers you like).
However, K8s is actually more about microservices. Each microservice can act like a tiny app on its own, but they are often inter-dependent and, especially at the beginning, it's often seen as not cost-effective to dedicate their own servers to them (along with the associate load balancing, redundant and cross-AZ, etc). And you might not even know what the scaling pain points for an app is, so this gives you a way to easily scale up without dedicating slightly expensive instances or support staff to running each cluster; your scale point is on the entire k8s cluster itself.
Even though that is ALL true, it's also true that k8s' sweet spot is actually pretty narrow, and many apps and teams probably won't benefit from it that much (or not at all and it actually ends up being a net negative, and that's not even talking about the much lower security isolation between containers compared to instances; yes, of course, k8s can schedule/orchestrate VMs as well, but no one really does that, unfortunately.)
But, it's always good resume fodder, and it's about the closest thing to a standard in the industry right now, since everyone has convinced themselves that the standard multi-AZ configuration of 2014 is just too expensive or complex to run compared to k8s, or something like that.
I had a different experience. Some years ago I wanted to set up a toy K8s cluster over an IPv6-only network. It was a total mess - documentation did not cover this case (at least I have not found it back then) and there was a lot of code to dig through to learn that it was not really supported back then as some code was hardcoded with AF_INET assumptions (I think it's all fixed nowadays). And maybe it's just me, but I really had much easier time navigating Linux kernel source than digging through K8s and CNI codebases.
This, together with a few very trivial crashes of "normal" non-toy clusters that I've seen (like two nodes suddenly failing to talk to each other, typically for simple textbook reasons like conntrack issues), resulted in an opinion "if something about this breaks, I have very limited ideas what to do, and it's a huge behemoth to learn". So I believe that simple things beat complex contraptions (assuming a simple system can do all you want it to do, of course!) in the long run because of the maintenance costs. Yeah, deploying K8s and running payloads is easy. Long-term maintenance - I'm not convinced that it can be easy, for a system of that scale.
I mean, I try to steer away from K8s until I find a use case for it, but I've heard that when K8s fails, a lot of people just tend to deploy a replacement and migrate all payloads there, because it's easier to do so than troubleshoot. (Could be just my bubble, of course.)
* Cert manager.
* External-dns.
* Monitoring stack (e.g. Grafana/Prometheus.)
* Overlay network.
* Integration with deployment tooling like ArgoCD or Spinnaker.
* Relatively easy to deploy anything that comes with a helm chart (your database or search engine or whatnot).
* Persistent volume/storage management.
* High availability.
It's also about using containers which mean there's a lot less to manage in hosts.
I'm a fan of k8s. There's a learning curve but there's a huge ecosystem and I also find the docs to be good.
But if you don't need any of it - don't use it! It is targeting a certain scale and beyond.
You can run monolithic apps with no downtime restarts quite easily with k8s using rollout restart policy which is very useful when applications take minutes to start.
I can bring up a service, connect it to a postgres/redis/minio instance, and do almost anything locally that I can do in the cloud. It's a massive help for iterating.
There is a learning curve, but you learn it and you can do so damn much so damn easily.
Now I have a small personal cluster with machines and vps's (on some regions I don't have enough deployments to justify an entire machine) with a distributed multi-site fs that's mostly as certified for workloads as any other cloud. CDN, GeoDNS, nameservers all handled within the cluster. Any machine can go offline while connectivity remains the same, minus the timeout requirement of 5 minutes of downed pods to be rescheduled for monolithic services.
Kubernetes also provides an amazing way to learn things like bgp, ipam and many other things via calico, metallb and whatever else you want to learn.
Every time I see one of these posts and the ensuing comments I always get a little bit of inverse imposter syndrome. All of these people saying "Unless you're at 10k users+ scale you don't need k8s". If you're running a personal project with a single-digit user count, then sure, but only purely out of a cost-to-performance metric would I say k8s is unreasonable. Any scale larger, however, and I struggle to reconcile this position with the reality that anything with a consistent user base should have zero-downtime deployments, load balancing, etc. Maybe I'm just incredibly OOTL, but when did these simple features to implement and essentially free from a cost standpoint become optional? Perhaps I'm just misunderstanding the argument, and the argument is that you should use a Fly or Vercel-esque platform that provides some of these benefits without needing to configure k8s. Still, the problem with this mindset is that vendor lock-in is a lot harder to correct once a platform is in production and being used consistently without prolonged downtime.
Personally, I would do early builds with Fly and once I saw a consistent userbase I'd switch to k8s for scale, but this is purely due to the cost of a minimal k8s instance (especially on GKE or EKS). This, in essence, allows scaling from ~0 to ~1M+ with the only bottleneck being DB scaling (if you're using a single DB like CloudSQL).
Still, I wish I could reconcile my personal disconnect with the majority of people here who regard k8s as overly complicated and unnecessary. Are there really that many shops out there who consider the advantages of k8s above them or are they just achieving the same result in a different manner?
One could certainly learn enough k8s in a weekend to deploy a simple cluster. Now I'm not recommending this for someone's company's production instance, due to the foot guns if improperly configured, but the argument of k8s being too complicated to learn seems unfounded.
/rant
k8s makes it easier to build over engineered architectures for applications that don’t need that level of complexity
So while you are correct that it is not actually that difficult to learn and implement K8S it’s also almost always completely unnecessary even at the largest scale
given that you can do the largest scale stuff without it and you should do most small scale stuff without it, the number of people for whom all of the risks and costs balancr out is much smaller than the amount that it has been promoted and pushed
And given the fact that orchestration layers are a critical part of infrastructure, handing over or changing the data environment relationship in a multilayer computing environment to such an extent is a non-trivial one-way door
I use it (specifically, the canned k3s distro) for running a handful of single-instance things like for example plex on my utility server.
Containers are a very nice UX for isolating apps from the host system, and k8s is a very nice UX for running things made out of containers. Sure it's designed for complex distributed apps with lots of separate pieces, but it still handles the degenerate case (single instance of a single container) just fine.
I find k8s an extremely nice platform to deploy simple things in that don't need any of the advanced features. All you do is package your programs as containers and write a minimal manifest and there you go. You need to learn a few new things, but the things you do not have to worry about that is a really great return.
Nomad is a good contender in that space but I think HashiCorp is letting it slowly become EOL and there are bascially zero Nomad-As-A-Service providers.
(or, if you read it in a frustrated voice, The F**ing Article.)
It's only useful for the degenerate "run lots of instances of webapp servers running slow interpreted languages" use case.
Trying to do anything else in it is madness.
And for the "webapp servers" use case they could have built something a thousand times simpler and more robust. Serving templated html ain't rocket science. (At least compared to e.g. running an OLAP database cluster.)
Oh, and just in case your first rebuttal is "having thousands of containers means you've already failed" - not everyone works in a mom n pop shop
Just because k8s is the only game in town doesn't mean it is technically any good.
As a technology it is a total shitshow.
Luckily, the problem it solves ("orchestrating" slow webapp containers) is not a problem most professionals care about.
Feature creep of k8s into domains it is utterly unsuitable for because devops wants a pay raise is a different issue.
I truly wish you were right, but maybe it's good job security for us professionals!
What aspects are you referring to?
>> is not a problem most professionals care about
professional as in True Scotsman?
No, I mean that Kubernetes solves a super narrow and specific problem that most developers do not need to solve.
The majority of folks, whether or not they admit it, probably do...
Must be some artifact of hosting on Azure, because I can't imagine any other reason to do something this contorted.
With all that said, while I have no "hate" for the stack, I still have no plans to migrate our container infrastructure to it now or in the foreseeable future. I say that precisely because I've seen the source, not in spite of it. The net ROI on subsuming that level of complexity for most application ecosystems just doesn't strike me as obvious.
* Its secrets management was terrible, and for awhile it stored them in plaintext in etcd. * The learning curve was real and that's dangerous as there were no "best practice" guides or lessons learned. There are lots of horror stories of upgrades gone wrong, bugs, etc. Complexity leaves a greater chance of misconfiguration, which can cause security or stability problems. * It was often redundant. If you're in the cloud, you already had load balancers, service discovery, etc. * Upgrades were dangerous and painful in its early days. * It initially had glaring third party tooling integration issues, which made monitoring or package management harder (and led to third party apps like Helm, etc).
A lot of these have been rectified, but a lot of us have been burned by the promise of a tool that google said was used internally, which was a bit of a lie as kubernetes was a rewrite of Borg.
Kubernetes is powerful, but you can do powerful in simple(r) ways, too. If it was truly "the most amazing" it would have been designed to be simple by default with as much complexity needed as everybody's deployments. It wasn't.
People who say "you don't need k8s" never say what you do need. K8s gives us a uniform interface that works for everything. We just have a few YAML files for each app and it just works. We can just chuck new things on there and don't even have to think about networking. Just add a Service and it's magically available with a name to everything in the cluster. I know how to do this stuff from scratch and I do not want to be doing it every single time.
- Platform build and management - App build and management
Getting a stable K8s cluster up and running is quite different to building and running apps on it. Obviously there is overlap in the knowledge required, but there is a world of difference between using a cloud based cluster over your own home made one.
We are a very small team and opted for cloud managed clusters, which really freed me up to concentrate on how to build and manage applications running on it.
- It can be complex depending on the third-party controllers and operators in use. If you're not anticipating how they're going to make your resources behave differently than the documentation examples suggest they will, it can be exhausting to trace down what's making them act that way.
- The cluster owners encounter forced software updates that seem to come at the most inopportune times. Yes, staying fresh and new is important, but we have other actual business goals we have to achieve at the same time and -- especially with the current cost-cutting climate -- care and feeding of K8s is never an organizational priority.
- A bunch of the controllers we relied on felt like alpha grade toy software. We went into each control plane update (see previous point) expecting some damn thing to break and require more time investment to get the cluster simply working like it was before.
- While we (cluster owners) begrudgingly updated, software teams that used the cluster absolutely did not. Countless support requests for broken deployments, which were all resolved by hand-holding the team through a Helm chart update that we advised them they'd need to do months earlier.
- It's really not cheaper than e.g. ECS, again, in my experience.
- Maybe this has/will change with time, but I really didn't see the "onboarding talent is easier because they already know it." They absolutely did not. If you're coming from a shop that used Istio/Argo and move to a Linkerd/Flux shop, congratulations, now there's a bunch to unlearn and relearn.
- K8s is the first environment where I palpably felt like we as an industry reached a point where there were so many layers and layers on top of abstractions of abstractions that it became un-debuggable in practice. This is points #1-3 coming together to manifest as weird latency spikes, scaling runaways, and oncall runbooks that were tantamount to "turn it off and back on."
Were some of these problems organizational? Almost certainly. But K8s had always been sold as this miracle technology that would relieve so many pain points that we would be better off than we had been. In my experience, it did not do that.
I remember when microservices architecture was the latest hot trend that came off the presses. Small and big firms were racing to redesign/reimplement apps. But most forgot they weren’t Google/Netflix/Facebook.
I remember end user experience ended up being _worse_ after the implementation. There was a saturation point where a single micro service called by all of the other micro services would cause complete system meltdown. There was also the case of an “accidental” dependency loop (S1 -> S2 -> S3 -> S1). Company didn’t have an easy way to trace logs across different services (way before distributed tracing was a thing). Turns out only a specific condition would trigger the dependency loop (maybe, 1 in 100 requests?).
Good times. Also, job safety.
I agree on how it's technically a waste of time to pursue fads, but it's also a huge PITA to have a platform that good engineers actively try to avoid, as their careers would stagnate (even as they themselves know that it's half a fad)
1. Has been at org for so long, management condones them flaunting the rules, like pushing straight to prod. Hates the "inefficiency" of open source platforms and purpose-built something "suitable for the company" by themselves, no documentation you have to ask them to fix issues because they don't accept code or suggestions from others. The DSL they developed is inconsistent and has no parser/linter.
absoutely start as simple as you can, but plan to move to a hosted kube something asap instead of writing your own base images, unless that's a differentiator for your company.
2. Fad-driven
I wonder why I don’t often see (1) critiqued on the basis of (2).
The classic dependency loop example that you thought will never encounter again for the rest of your life after OS class
As someone using Kafka, I'd like to know what the (good) alternatives are if you have suggestions.
Where I'm at, most of Kafka usage adds nothing of note and could be replaced with a rest service. It sounds good that Kafka makes everything execute in order, but honestly just making requests block does the same thing.
At least then I could autoscale, which Kafka prevents.
I found it easy to get up and running, even as a RAFT cluster, but I have not tried to use JetStream mode heavily yet.
There is really nothing wrong with a large vertically scaled up SQL server. You need to be either really really large large scale - or really really UNSKILLED in sql as to keep your relational model and working set in SQL so bad that you reach its limits
Sadly that's the case at my current job. Zero thought put into table design, zero effort into even formatting our stored procedures in a remotely readable way, zero attempts to cache data on the application side even when it's glaringly obvious. We actually brought in a consultant to diagnose our SQL Server performance issues (I'm sure we paid a small fortune for that) and the DB team and all of the other higher ups capable of actually enforcing change rejected every last one of his suggestions.
Large or small, it's always going to be a single point of failure. All hardware is fallible, especially if we're talking about a commodity box rather than a mainframe or something.
It works, with a lower cognitive burden than that of horizontally scaling.
For the loading concern (i.e. is this enough to handle the load):
For most businesses, being able to serve 20k concurrent requests is way more than they need anyway: an internal app used by 500k users typically has fewer than 20k concurrent requests in flight at peak.
A cheap VPS running PostgreSQL can easily handle that.[1]
For the "if something breaks" concern:
Each "fault-tolerance" criteria added adds some cost. At some point the cost of being resistant to errors exceeds the cost of downtime. The mechanisms to reduce downtime when the single large SQL server shits the bed (failovers, RO followers, whatever) can reduce that downtime to mere minutes.
What is the benefit to removing 3 minutes of downtime? $100? $1k? $100k? $1m? The business will have to decide what those 3 minutes are worth, and if that worth exceeds the cost of using something other than a single large SQL server.
Until and unless you reach the load and downtime-cost of Google, Amazon, Twitter, FB, Netflix, etc, you're simply prematurely optimising for a scenario that, even in the businesses best-case projections, might never exist.
The best thing to do, TBH, is ask the business for their best-case projections and build to handle 90% of that.
[1] An expensive VPS running PostgreSQL can handle a lot more than you think.
Not to forget: those costs are not just in money and time, but also in complexity. And added complexity comes with its own downtime risks. It's not that uncommon for systems to go down due to problems with mechanisms or components that would not exist in a simpler, "not fault tolerant" system.
That's still a business decision.
Customers don't vote with their feet based on what tech stack the business chose, they vote based on a range of other factors, few, if none, of which are related to 3m of downtime.
There are little to no services I know off that would lose customers over 3m of downtime per week.
IOW, 3m of downtime is mostly an imaginary problem.
Services that people might leave, because of downtime are for example a git hoster, or a password manager. When people cannot push their commits and this happens multiple times, they may leave for another git hoster. I have seen this very example when gitlab was less stable and often unreachable for a few minutes. When people need some credentials, but cannot reach their online password manager, they cannot work. They cannot trust that service to be available in critical moments. Not being able to access your credentials leaves a very bad impression. Some will look for more reliable ways of storing their credentials.
> The business can try to decide what those 3min are worth, but ultimately the customers vote by either staying or leaving that service.
What do you think the business is doing when it evaluates what 3 minutes are worth?
I don't understand, what people are arguing about here. Are we really arguing about customers making their own choice? Since that is all I stated. The business can jump up and down all it wants, if the customers decide to leave. Is that not very clear?
I think the point is that, for a few minutes of downtime, businesses lose so little customers that it's not worth avoiding that downtime.
Just now, we had a 5m period where disney+ stopped responding. We aren't going to cut off our toddler from peppa big and bluey for 5m of downtime per day, nevermind per week.
You appeared to be under the impression that 3m downtime/week is enough to make people leave. This is simply not true, especially for internet services where the users are conditioned to simply wait.
You saying “the customer can decide to leave” is not a counter to this at all. It’s just a weird way of saying what everyone else is but framing it as a counter argument to what is being said. Which it simply isn’t.
I've never, in my life, seen an error in SQL Server related to SQL Server. It's always been me, the app code developer.
Now, to be fair, the server itself or the hardware CAN fail. But having active/passive database configurations is simple, tried and tested.
I wish we would use something like Debian and take advantage of tech like systemd. But alas, we're still using COM and Windows Services and we still need to remote desktop in and click around on random GUIs to get stuff to work.
Luckily, SQL Server itself is very stable and reliable. But even SQL Server runs on Linux.
This is a very simple distinction and I am not sure why is it not understood
For some reason people design public apps the same as internal apps
The largest companies employ circa 1 million people - that’s Walmart, Amazon, etc. most giants, like Shell, etc. companies have ~ 100k tops. That can be handled by 1 beefy server.
Successful consumer facing apps are hundred millions to billion. It’s 3 orders of magnitude difference
I have seen a company with 5k employees invest into mega-scalable microservice event driven architecture and I was was thinking - I hope they realise what they are doing and it’s just CV-driven development
Agreed, but the cost/benefit should be the same at every level in the stack. I get it when people build a single-server system. I'm just baffled that so many people insist they need a load balancer and multiple instances of their application so that there's no single point of failure (which is not free by any means), but then run them all off of a single SQL server.
It's less true today when redundancy is baked into SaS products (like AWS Aurora, where even if you have a single database instance, it's easy to spin up a replacement one if the hardware on the one running fails).
One of the most stable system archictectures I've built was on Kafka AND it was running with microservices managed by teams across multiple geographies and time zones. It was one of the most reliable systems in the bank. There are situations where it isn't appropriate, which can be said for most tech e.g( K8S vs ECS vs Nomad vs bare metal)
Every system has failure characteristics. Kafka's is defined as Consistent and Available and your system architecture needs to take that into consideration. Also the transactionality of tasks across multiple services and process boundaries is important.
Let's not pretend that kubernetes (or the tech of the day) is at fault while completely ignoring the complex architectural considerations that are being juggled
From my experience most microservices aren’t engineered to handle back pressure. If there is a sudden upsurge in traffic or data the Kafka cluster is expected to absorb all of the throughput. If the cluster starts having IO issues then literally everything in your “distributed” application is now slowly failing until the consumers/brokers can catch up.
Yep - saw that at a company recently - something in AWS was running a little slower than usual which cascaded to cause massive failures. Dozens of people were trying to get to the bottom of it, it mysteriously fixed itself and no one could offer any good explanation.
Is that the same as what Elixir/Erlang call supervision trees?
But to the point, if you're going to build a distributed system, you need tools to track problems across the distributed system that also works across teams. A poorly performing service could be caused by up/downstream components and doing that without some kind of tracing is hard even if your stack is linear.
The same is true for a giant monolithic app, but the sophisticated tools are just different.
There's very little reason to distribute application code. It's very, very rare that the limiting factor in an application is compute. Typically, it's the data layer, which you can change independently of your application.
But I think very few problems fit into that archetype. Instead, people build distributed systems for reliability and integrity. But it's overkill, because you bring all the baggage and complexity of distributed computing. This area is incredibly difficult. I view it similar to parallelism. If you can avoid it for your problem, then avoid it. If you really can't, then take a less complex approach. There's no reason to jump to "scale to X threads and every thread is unaware of where it's running" type solutions, because those are complex.
It's very rare to meet a developer who has even the vaguest notion of what an RPC call costs in terms of microseconds.
Fewer still that know about issues such as head-of-line blocking, the effects of load balancer modes such as hash versus round-robin, or the CPU overheads of protocol ser/des.
If you have an architecture that involves about 5 hops combined with sidecars, envoys, reverse proxies, and multiple zones you're almost certainly spending 99% to 99.9% of the wall clock time just waiting. The useful compute time can rapidly start to approach zero.
This is how you end up with apps like Jira taking a solid minute to show an empty form.
I think Web apps have been around so long now people have forgotten how unresponsive things are vs old 2 tier stuff.
I tested the Jira cloud service with a new, blank account. Zero data, zero customisations, zero security rules. Empty.
Almost all basic operations took tens of seconds to run, even when run repeatedly to warm up any internal caches. Opening a new issue ticket form was especially bad, taking nearly a minute.
Other Atlassian excuses included: corporate web proxy servers (I have none), slow Internet (gigabit fibre), slow PCs (gaming laptop on "high performance" settings), browser security plugins (none), etc...
At that point, something mustve been wrong with your instance. I'd never call jira fast, but the new ticked dialog on a unconfigured instance opens within <10s (which is absolutely horrendous performance to be clear. Anything more then 200-500ms is.)
it's open within 1-2 sec (which is still bad performance, objectively speaking. It's an empty instance after all and already >1s )
It's still sluggish compared to a desktop app from the 1990s, but it's much faster than just a couple of years ago.
To be fair, small time units are difficult to internalize. Just look at what happens when someone finds out that it takes tens of nanoseconds to call a C function in Go (gc). They regularly conclude that it's completely unusable, and not just in a tight loop with an unfathomable number of calls, but even for a single call in their program that runs once per day. You can flat out tell another developer exactly how many microseconds the RPC is going to add and they still aren't apt to get it.
It is not rare to find developers who understand that RPC has a higher cost than a local function, though, and with enough understanding of that to know that there could be a problem if overused. Where they often fall down, however, is when the tools and frameworks try to hide the complexity by making RPC look like a local function. It then becomes easy to miss that there is additional overheard to consider. Make the complexity explicit and you won't find many developers oblivious to it.
I'm not too familiar with Go but my default assumption is that it's just used as a convenient excuse to avoid learning how to do FFI.
Nah. Jira was horrible and slow long before the microservice trend.
In fact, I'd argue that one of the biggest annoyances I had in last job was that I actually missed JIRA, or rather it's glorious issue list view with its keyboard shortcuts.
So really eventually we'll all be optimizing around network interfaces as the bottleneck.
Of course, it was because the client app recently went over the 100MB JS budget. Which they decided to make because the last time that happened, customers abroad reported seeing “white screens”. International conversion dropped sharply not long after that.
It’s pretty silly. So ya, good times indeed. Time to learn k8s.
They already has their architecture, largely, and just moved it over to K8s.
They even mention they aren’t a micro service company.
to illustrate performance cost i usually ask people what ping they have say from one component/pod/service to another and to compare that value to what ping they'd get between 2 linux boxes sitting on that gorgeous 10Gb/40Gb/100Gb or even 1000Gb network that they are running their modern microservices architecture over.
> More broadly, we’re not a microservices company, and we don’t plan to become one
Is this generally correct? How well is this term defined?
If yes, then I'm not surprised this type of design has become a target for frustration. State is smeared across system, which implies a lot of messaging and arbitrary connections between services.
That type of design is useful if you are an application platform (or similar) where you have no say in what the individual entities are, and actually have no idea what they will be.
But if you have the birds-eye view and implement all of it, then why would you do that?
And the "test cases" porn! Goodness! People want "coverage" and that's it. It's a box to tick, irrespective of whether those test cases and they way those are written are actually meaningful in anyway. There are more like "let's have dependency injection everywhere" charade.
grug wonder why big brain take hardest problem, factoring system correctly, and introduce network call too
seem very confusing to grug
https://grugbrain.dev/#grug-on-microservicesSo they didn't do microservices correctly. Big surprise.
There are certainly ways to mitigate the SPOFiness of each of those cases, but that doesn’t make having them an antipattern.
Personally, i think, it does not have to be hundreds of microservices, basically each function a service. But i see it more as the web in the internet. Things are sometimes not reachable or overloaded. I think that is normal life.
Kubernetes does not mean microservices, it does not mean containerization and isolation, hell, it doesn't even mean service discovery most of the time.
The default smallest kubernetes installation provides you two things: kubelet (the scheduling agent) and kubeapi.
What do these two allow you to do? KubeApi provides an API to interact with kubelet instances by telling them to do with manifests.
That's all, that's all kubernetes is, just a dumb agent with some default bootstrap behavior that allows you to interact with a backend database.
Now, let's get into kubernetes default extensions:
- CoreDNS - linking service names to service addresses.
- KubeProxy - routing traffic from host to services.
- CNI(many options) - Networking between various service resources.
After that, kubernetes is whatever you want it to be. It can be something that you can use to spawn few test databases. Deploy an entire production-certified clustered databases. A full distributed fs with automatic device discovery. Deploy backend services if you want to take advantage of service discovery, autoscaling and networking. Or it can be something as small as deploying monitoring (such as node-exporter) to every instance.
And as a bonus, it allows you to do it from the comfort of your own local computer.
This article says that figma migrated necessary services to kubernetes to improve developer experience and clearly said that things that don't need to be kubernetes aren't. For all we know they still run their services in raw instances and only use kubernetes for their storage and databases. And to add to all of that, kubernetes doesn't care where it runs, which is a great way to increase competition between cloud providers lowering the costs for all.
This ties into a funny example: k8s manages my vm's via kubevirt, those then have a minimal k8s version installed that runs my jobs. The implementation simply mounts the extracted image to a virtual fs and executes it there, then deletes the file system.
But that's assuming you're running containerd or something similar. There are dozens of k8s implementations some as light as only providing you with manifests which then you have external schedulers which subscribe (called controllers) that execute on these manifests.
If you don't have a distributed system then personally I think k8s makes no sense.
I agree that for a handful of pet-servers for a team with more existing Linux experience than k8s experience, this is a better starting point, because of the shorter learning curve. Just not kid ourselves that the end product has any less complexity, it's only a different skill set.
... write a unit file and put it in your CI?
> How do you run a container in systemd?
systemd-nspawn config, which you put in your aforementioned unit files.
> perhaps Ansible and docker-compose
Definitely don't need docker compose and, in my opinion, don't need Ansible. It's trivial to deploy systemd config and it's trivial to automate.
I'm not kidding myself - the complexity is significantly lower. But it only works if you're deploying to one, or two, machines. This won't make a distributed system and I acknowledge that.
Not to mention I'm only scraping the surface of what systemd can do here. Containers and automating services is just part of it. There's also remoting logging, monitoring and email alerts, periodic health checks.
Yeah, at those low growth companies, you have unlimited resources /s