(Hard-won bit of experience: k8s + Redis really don't like each-other if Redis 1. is configured to load from disk, and 2. your memory limit for the Redis container is somewhat-tightly bounded. At least from the k8s controller's perspective, Redis apparently uses ~400% of its steady-state memory while reading the AOF tail of an RDB file — getting the container stuck in an OOM-kill loop until you come along and temporarily de-bound its memory.)
However, we're considering switching back to k8s for stateful components, with a different approach: allocating single-node node-pools with taints that map 1:1 to each stateful component, effectively making these more like "k8s-managed VMs" than "k8s-managed containers." The point would be to get away from the need to manage the VMs ourselves, giving them over to GKE, while still retaining the assumptions of VM isolation (e.g. not having/needing memory limits, because the single pod is the only tenant of the VM anyway.)
I use Kubernetes to host something like 6 of our own services, and it excels at that and is fairly simple. Databases and other things use different services.
Do you really think they have ALL the operations coded properly for all operational conditions?
Stateless is so much easier for operations. It either is running, or not, A/B upgrades, yada yada.
Stateful has backups, restores, outages, patching, migrations, corruptions, fixes. For distributed systems it gets even hairier. Do they support your distribution? What if you have a multicloud or blended internal/cloud? Does the operator provide a turnkey restore from backup, and are you testing it?
Operators shouldn't be viewed as replacements for stateful operations knowledge, but it's probably what they'll be used for.
If you're using a large scale distributed stateful/database system, you need an operations team to support it, or pay for that.
We've just tuned it for a java-based app which was also stuck in OOMkill hell, and this has completely resolved the situation (MALLOC_ARENA_MAX=2).
That's a problem we're hoping to help solve, where you define your application and it's dependencies, and we help run it in the right way to leverage managed cloud services across environments without passing that headache on!
https://cloud.google.com/compute/docs/instance-groups#manage...
ES handles node restarts or upgrades pretty gracefully though. I'd imagine for databases or "non-clustered" things you'd have to consider GKE's aggresive upgrade schedule. We use CloudSQL for some databases but our larger ones are still on GCE because we get more control of replication, CDC, and can use tools like proxysql to reduce downtime.
Devil's advocate but isn't having to maintain VMs (and then software deployed to those VMs) and k8s YAML/charts/whatever more "maintenance burden" than just one or the other?
Just something that takes in API requests (or a cron-like scheduled job) and makes other API calls/does plumbing?
We used managed services for our stateful stuff which significantly eased the operational burden there. Might be a different story if we looked at doing the absolute minimum cost optimization. However, at least for us, the extra cost of managed services is worth the price.
The yaml tends to be a "one and done" sort of thing. We touch it MAYBE once every 2 months if that.
Ephemeral Redis is well-suited to k8s — you can treat it as just a sidecar to your app layer deployment. Durable Redis is not. But sometimes the transition can sneak up on you.
Isn't this just moving the problem from per-pod resource constraints to per-VM resource constraints?
nodeAffinity:
required:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/hostname
operator: In
values:
- <hostname>
this ensures that your workload will be rescheduled on the matching nodewhat’s the difference between doing this as single node pools vs pod constraints like anti affinity?
It can also make monitoring resource usage a little easier since you can just monitor at the node level
As for K8 managed VMs - thats a good idea. Resource management wise, security wise etc.
On the other hand, once everyone on the team has experience building such a system from scratch, then deploying k8s and using it somehow becomes straightforward.
It's almost as if we need to learn how a tool works before being able to use it effectively.
Anyways, what we (actually didn't) replace it with:
- Don't let your devs learn about k8s on the job.
- Let them run side-projects on your internal cluster.
- Give them a small allowance to run their stuff on your network and learn how to do that safely.
- Give your devs time to code review each other's internally-hosted side-projects-that-use-k8s.
- Reap the benefits of a team that has learnt the ins and out of k8s without messing up your products.If somebody else on the team already knows k8s. The problem is a lot of places just give devs admin access and let them go hog wild. If devs don't know k8s they can't make significant changes without waiting for the one guy who knows k8s to do it.
Give a man a k8s, he will deploy for a day. Teach a man to k8s, he will deploy for at least 5 years while the hype cycle continues.
1. Are motivated by the sort of curiosity that would frustrate them if they were blocked from knowing about k8s
2. Are motivated by the sense of responsibility that would unnerve them if they didn't understand 1 abstraction layer beneath their work.
?
If you want your Devs to learn kubernetes you should pay them for doing it.
If you can't, hire a contractor with the expertise you need.
It is difficult to tell at a glance whether an engineer is qualified to effectively use a tool. Letting them self-train by working on self-projects in isolation compounds this effect.
The goal of this exercise is to give time and space to your devs to practice in a safe environment, while allowing them to push, deploy and review projects internally as if they were core products, so that other SMEs are allowed to spend some time every week reviewing those projects for smells and issues before those ideas make it into a core product.
Fortunately, k8s can act as a very secure sandbox when it's configured properly, so you'll know how to mitigate such a disaster once your company has trained its engineers on how to use the tool effectively.
It's not as common as some would have you believe but it is real. A teammate spent his 20% on underwater topography on Google Earth and another spent his on the glider thing.
Cattle not pets after all.
Note: this isn't everyone's end game but I suspect it's realistic for a lot of people.
I would like to go back to cleanly divided, architected IaaS and ansible. It was fast, extremely reliable, cheaper to run, had a much lower cognitive load and a million less footguns. What's more important possibly is not everything can be wedged into containers cleanly despite the promises.
This is the mental abstraction I've been operating with for over a decade now.
All of our products are monolithic binaries that can be installed on bare-ass windows or linux machines. For all intents & purposes, basic AWS/Azure/et. al. VM hosting is our containerization strategy. We just pushed the tricky bits down into our software.
95% of our pain is resolved by using a modern .NET stack and leaning hard on their Self-Contained Deployment model. Our software has zero external dependencies at deploy time, so there isn't much to orchestrate. Anything that talks to a 3rd party system is managed purely via configuration in our software.
I'm working with a product now that's made their k8s deployment the standard and all it's done is create bigger issues. Ops got behind on Strimzi and so we got stuck on 1.21 because we couldn't upgrade due to being locked to the Strimzi version. This caused issues because of log4j and we ran into a wall quickly with customers on GCP as soon as 1.22 ended up as GA. Honestly I'm not sure we're getting much, if any, overhead advantage since I feel like the app has become bloated due to container creep.
That and supporting 4 different ways to provision storage across customers on every cloud / on-prem is a nightmare. Customer environment installed applications on k8s is a nightmare today.
A VM provides better isolation than a container as it has a separate kernelspace. Hence the DevOps mantra "containers don't contain".
You could look at containerization without K8S (podman or docker) especially if you use python and don’t want to mess with the Linux native python installation.
It's original purpose wasn't to do elastic scaling or anything like that - it was to binpack workloads onto a set of nodes, and not everyone has Silly Valley money to pay Silly Valley prices (especially when one's currency is weak against dollar)
This is arguably still its primary purpose, and all the rest of its features are ancillary and only exist for the sake of operational convenience.
We ended up writing our own control plane that uses NATS as a message bus. We are in the process of open sourcing it here: https://github.com/drifting-in-space/spawner
Just curious if you could elaborate here? I work with k8s on docker, and we're also going to be spinning up ephemeral containers (and most of the other things you say) with jupyter notebooks. We're all in on k8s, but since you might be ahead of me, just wondering what hurdles you have faced?
Our big problem was fetching containers took too long since we have kitchen sink containers that are like 10 GB (!) each. They seem to spin up pretty fast though if the image is already pulled. I've worked on a service that lives in the k8s cluster to pull images to make sure they are fresh (https://github.com/lsst-sqre/cachemachine) but curious if you are talking about that or the networking?
From what it looks like in your repo it might be that you need to do session timing (like ms) response time from a browser?
It wasn't really one thing with Kubernetes that was slow, but that the more we tried to optimize it the less of core Kubernetes we were using and so the less value we were getting for the complexity tax we were paying. The image pulling you mention is a good example of that; having pre-pulled images is a big factor, but we have too many images to push every image to every node, instead we'd like the scheduler to be aware of which node has which image. We could do that with node affinity, but what we'd end up building would be more work than if we wrote our own scheduler to support it from day one.
> From what it looks like in your repo it might be that you need to do session timing (like ms) response time from a browser?
Our goal is subsecond container starts. We're not there yet, and might not get there with Docker, but we have a POC that is there with WebAssembly-based workloads. Too bad those are rare :)
(By the way, I'm always happy to chat about this stuff, my email is in my profile)
The kubernetes scheduler should be aware of which node has which image, that is why the Node object has the status.images field: https://kubernetes.io/docs/reference/generated/kubernetes-ap....
It turned out to be somewhat tricky, because it increased the size of the Node object, and colocating node heartbeats onto the same object meant that a bigger object was changing relatively often. But that was addressed by moving heartbeats to a different object: https://github.com/kubernetes/enhancements/issues/589
It doesn't get all the way to what we want, but it could be used to build a piece of it.
[1] Gigabytes in milliseconds: Bringing container support to AWS Lambda without adding latency. https://www.youtube.com/watch?v=A-7j0QlGwFk
https://cloud.google.com/kubernetes-engine/docs/how-to/image...
> the more we tried to optimize it the less of core Kubernetes we were using and so the less value we were getting for the complexity tax we were paying
Since we were headed down that path, we took a step back and asked what we were really getting out of Kubernetes, and most of it was things that were orthogonal to our intended use case. The way Kubernetes is architected around control loops works great for its intended use case, but we wanted a more event-driven system.
Also there's that project called Nydus that allows starting up big containers way faster. IIRC, starts the container before pulling the whole image, and begins to pull data as needed from the registry.
https://github.com/senthilrch/kube-fledged
https://github.com/dragonflyoss/Dragonfly2
https://github.com/containerd/stargz-snapshotter/blob/main/d...
One question: have you had to do any custom container builds on demand, and if so, have you had to deal with large "kitchen sink" containers (e.g. a Python base image with a few larger packages installed from PyPI, plus some system packages like Postgres client)? We would run up against extremely long build image times using tools like kaniko, and caching would typically have only a limited benefit.
I was experimenting using Nix to maybe solve some of these problems, but never got far enough to run a speed test, and then left the job before finishing. But it seems to me some sort of algorithm like Nixery uses (https://nixery.dev) to generate cacheable layers with completely repeatable builds and nothing extraneous would help.
Maybe that's not a problem you had to solve, but if it is, I'd love your thoughts.
Is that not how this works?
I wrote the student vm system for udacity, and I spun up student vms before they needed them, with some last mile loading to finalize the files they need. The student VMs were not using k8s, although a small piece of the infrastructure did.
I worried most about untrusted users working in a complex environment with the ability to harm the experience of other users, and just used GCE.
For me, boot time was < 5 minutes, so if you can predict the next five minutes of demand, you can boot those machines early. If you are wrong you will pay extra or someone will wait extra time, but still less than the full boot time. Generally it takes less than 10 seconds to access a vm with your coursework on in.
One of my ideas lately has been to upgrade FaaS to a full on server after a set amount of traffic. Or said differently, a dedicated server spin up that serves the same app as callable functions ala scalable RPC and upgrade to a dedicated instance composed of said functions. The best of both worlds.
Combine the scale to zero of Serverless combined with the scalability and capacity of a dedicated server.
But on other hand, I can't quite figure out why something would prevent, you, yourself, from running the service that hosts the VMs that hosts the containers on demand on Kubernetes.
> But on other hand, I can't quite figure out why something would prevent, you, yourself, from running the service that hosts the VMs that hosts the containers on demand on Kubernetes.
I'm not sure I understand this part, I guess we could use Kubernetes operators to scale up the underlying compute resources and manage the containers ourselves? This adds a lot of complexity for our use case.
When using K8, if you use the most basic K8 features and concepts, things generally work out pretty ok.
From the Operations side, Kubernetes is scary. It's easy to screw things up and you can definitely run into problems. I understand why folks who work mostly on that side of the house are put off by the complexity of Kubernetes.
However, from the application side of things, our developers have been THRILLED with Kubernetes. For most developers my company provides a nice paved road experience with minimal customization required. For advanced use cases, we allow developers to use the Kubernetes API (along ArgoCD + GateKeeper policies) as a break glass type of approach. Istio gives the infra team the ability to easily move services between clusters and make policy changes easily. It also allows us to make use of Knative, although I think the Istio requirement is no longer there.
That said, you should be using managed Kubernetes wherever possible and not running your own clusters. That's where trouble lurks.
It makes it that much easier to actually use the cluster rather than mess with endless configuration tooling. Is it the best engineered tool? Probably not. But it's the one that works best for us.
The result of the migration was that there is little underlying infrastructure to maintain, and ongoing operational costs were lowered by 50% year over year. The CTO and I liked the setup so much, we started converting another large client of theirs. I followed up with them at the beginning of 2022 to see how things were going, and they still love it. There is so little maintenance, and now they have more time to focus on what they do best–Software!
Other options on the horizon that I'm testing include utilizing AWS Copilot with ECS/Fargate, and/or Copilot with Amazon App Runner.
I still wonder from time to time if I am missing something not going Kubernetes.
If a team were to start with no legacy and no complexity and there isn't going to be multi-team/multi-owner/shared-services I could see them using something else. But that applies to anything.
Wish there was some better docs out there, not sure if I could handle writing one from scratch :/
These days I'm a huge fan of CDK and Pipelines style deployments. I prefer to treat my compute layer as a swappable component which I'll change as and when I need to. I tend to lean towards serverless offerings which take care of the internal scaling details if I can while still giving me a traditional "instance", and if I can't then I'll go for the next best managed offering.
I've yet to see an example where internal tooling doesn't become a mess over time, and K8S requires a ton of work to keep things sensible.
In the case of my org, we optimized for the features we thought were valuable and amortized that effort over time. Notably this was early in k8s history (2014/2015), but the fruits of those efforts have aged well so far (8 years or so). Small code footprint to cover service discovery, cert provisioning, release orchestration, and configuration management. The whole devops stack is less than 3k SLOC. Service ecosystem is ~150 distributed systems, roughly about 5 million SLOC, running on just over 1k servers on AWS.
I think if the aim is not to completely replace what k8s does, but to cherry pick the features that give you some pareto distribution of value, sometimes its worth it to build in-house. Nothing wrong of course with going with k8s for many orgs, but in our case we didn't have to reinvent the whole wheel to live without it.
You can actually get a couple of pretty beefy bare metal boxes for that budget. Or a couple of more modest ones for app servers plus a nice big RDS instance with all the trimmings. Based on past experience, that’ll get you to a few hundred rps for even a fairly complicated, poorly-tuned Rails or PHP app; your well-factored Go API server should handle 10x that pretty easily.
You might have to write some Bash or systemd unit files instead of a bunch of YAML, which may or may not bug you. I find shell easier to understand and debug than YAML-based scripting but YMMV.
2. Remote development. With k8s we can develop right out of the cluster, ECS has no comparative.
3. Installing OSS software. K8s has loads of supported packages for OSS tooling.
I remember managing hundreds of virtual machines in datacenters & cloud, using Ansible and a myriad of other tooling.
It's nice when you're at a small scale and you don't have a lot of people making changes, but over time as it grows the pain grows with it unless you've enforced a consistent cattle model.
The longer VMs live with custom changes/code and updates over time the more brittle they can become. Part of the cattle model is so that you can recreate/rebuild when changing code so things stay consistent. The drift from infrastructure as code can be scary otherwise.
With the cattle model you need to have pipelines in place to build new VM images for infrastructure updates (packer etc), have multiple APIs to hit (easier in cloud) to upload images and serve them in a non damaging way. (HA deployments/rollouts/dealing with load balancers) It's certainly a non-trivial amount of work.
With Kubernetes, a lot of this tooling comes out of the box. You've got autoscaling, load balancing, health-checks, limits/requests, failure mitigation, service mesh options. On top of that it's served in a strict semi-consistent way. Good luck replicating that with virtual machines without a lot of tooling and effort.
If you can learn the Kubernetes tooling it can do a lot for you. However I agree that not all setups need it, a lot of times small setups never grow and that's ok a few virtual machines aren't that big of a deal.
We still use virtual machines for workloads that aren't container friendly, and to be honest these days I abhor it, even with pipelines in place.
Honestly kubernetes is not harder than dealing with this. It's keeping you back in the land of default google-able problems longer as weird tweaks and unique configs aren't piling up to make esoteric issues.
Replaced with Linux servers and SSH.
Have done a lot of work with k8s in the past. Not the right tool for my startup.
We use systemd to manage services.
We use Ansible to set up servers.
Our infrastructure spans AWS, Google Cloud, and servers in a datacenter.
They are a better fit because they are much easier to manage and it's much easier to debug issues when something goes wrong. We are a small team and we hope to stay that way. But our operational responsibilities are growing significantly. The extra cognitive overhead of working with technologies like kubernetes would prevent us from scaling up our effort the way that we want to.
It's hard to answer your question in detail outside of a very long essay.
Linux namespaces are a brilliant idea.
With larger teams, I have written and maintained custom k8s operators in production. It was a great fit for the problems we had at that scale of developers.
A terrible fit for the problems my current team has.
There's still use-cases where k8s wins; but nomad handles state a bit better and is easier to reason about from scratch.
1) I don't really want to manage the installation but there aren't any(?) cloud hosts for nomad that I can see. 2) It doesn't seem as widely used so community support seems thin. There aren't many blog posts about good patterns with it etc, and I'd worry that we'd get stuck and end up reverting back to k8s.
Point 2 is debatable. Lots of people nowadays put Kubernetes on their resume but that doesn't mean they are great architects or technicians, yet a good part of running production on Kubernetes is doing it right.
You'll see much fewer people with Nomad on their resume, but on the other hand you know they're not here for the buzz, they're usually more experienced and know what they're talking about.
We wrote about our decision to switch here: https://www.koyeb.com/blog/the-koyeb-serverless-engine-from-...
Nomad, Consul and Vault interoperate extremely well and are mostly pleasant to use, but I found myself missing the rest of the ecosystem pretty quickly, especially around ingresses, and I think they made the wrong decision on the networking model compared to Kubernetes.
That said, I haven't played with Consul Connect, the Consul+Envoy service mesh, yet. That might address a lot of the problems. But fundamentally I can't help but think that Nomad and Kubernetes both made a run of it and Kubernetes came out the winner of mindshare and ecosystem.
Can you elaborate on how Nomad handles state differently than K8S and what makes it better?
- Taking random .yml configs from The InternetTM to install an Nginx Ingress with automatic LetsEncrypt certs felt not-exactly-great. It's no better than piping curl to bash, except the potential impact is not that your computer is dead, but the entirety of prod goes down.
- Because of this, upgrades of Kubernetes are a pain. The DigitalOcean admin panel will complain about problems in 'our' configs, that aren't actually OUR configs. We don't know how to fix that, or if ignoring the warnings and upgrading will break our production apps.
- Upgrades of Kubernetes itself aren't actually zero downtime, and we couldn't figure out how to do that (even after investing a significant amount of research time).
- We were using only a tiny subset of the functionality in Kubernetes. Specifically we wanted high-availability application servers (2+ pods in parallel) with zero-downtime deployments, connecting to a DO managed PostgreSQL instance, with a webserver that does SSL-termination in front of it.
- Setting up deployments from a GitLab CI/CD pipeline was pretty hard, and it turned out the functionality for managing a Kubernetes cluster from GitLab was not really done with our use case in mind (I think?).
- It would be bad enough if DigitalOcean shit the bed, but the biggest problem was that we couldn't reliably recognize if something was a problem caused by us, or by DO. Try explaining that one to your customers.
Summarizing: it was just too complex and fragile, even once you wrap your head around what the hell a Pod, a Deployment, an Ingress and Ingress Controller, and all of the other Kubernetes lingo actually means. I suspect you need a dedicated infra person who knows their stuff to make this work, so it could very well make sense for larger companies, but for our situation it was overkill.We were not intellectually in control of this setup, and I do not feel comfortable running production workloads (systems used by 20k high-school students, mission-critical applications used by logistical companies) on something we couldn't quite grasp.
We went to a much simpler setup on Fly.io, and have been happy since. It's a shame they seem to be too young of a company to really be super reliable, but I suspect this is only a matter of time. In terms of feature set, it's all we need.
I can pretty confidently say, that's not K8s, that's Digital Ocean. On AWS, we ran the EKS infrastructure (which was not simple) with basically half a dev's time for years. It was only when it started to scale to millions of users that we needed to build a team to support it. It was still a much smaller team than the one that supported the ECS product (two devops).
I was mostly managing and not coding by the time Kubernetes was in our stack, so while I'm very familiar with infrastructure in general (and I know ECS inside and out unfortunately), I hadn't used Kubernetes directly much before I build this DO infrastructure. But I got it up in a week and though DO is a nightmare, k8s is an absolute joy as a DevOps. Holy shit it's perfect. It does exactly what it needs to, with exactly the right abstractions, with perfectly reasonable defaults.
The reality is that infrastructure work is just that complicated.
You wouldn't try to have a team of front end engineers build your rest backend. It's not reasonable to expect javascript engineers to know how to build and operate an infrastructure - at least not with out dedicating themselves to learning the tooling and space full time for a while. Think of it from the perspective of a frontend engineer learning Python and Django to build out a rest backend, and then multiply the complexity by 4. That's just infrastructure regardless of what you're using.
That said, if something like Fly.io can fit your needs, that's great! I haven't used them so I can't speak to them directly, but I know that with Heroku, the trade off was cost and, eventually, being limited in what you could build. Eventually you would need to build something that just couldn't be built with Heroku. A quick glance at Fly, the pricing looks reasonable, but I'm guessing the build limits will still apply.
> The reality is that infrastructure work is just that complicated.
Yes, if you need the flexibility of running anything in any setup. What we really wanted was 'yeet a docker image with a web server in it + env vars at some magic beast that'll run it for me, slap an SSL-cert on it, and make sure it's always online'. We tried to replicate this with Kubernetes, so we got the full complexity of k8s unloaded upon us.
Heroku was what we really wanted, but it was always too expensive. Fly.io strikes a good balance here, the defaults are sane, it's still flexible enough for other services, and it's relatively cheap (spend is similar to DO K8s).
> You wouldn't try to have a team of front end engineers build your rest backend.
Well, yes and no. I wouldn't expect frontend engineers to know the ins and outs of everything backend, but to build on your metaphore a bit further: Setting up a basic Node backend with express serving static files shouldn't take multiple weeks, even for a frontend engineer. I feel like I was trying to do the infra equivalent of that, and it did take me forever.
> A quick glance at Fly, the pricing looks reasonable, but I'm guessing the build limits will still apply
The build limits could be an issue but really isn't for us right now. It's fairly easy to build locally though (in our case: in our GitLab CI/CD runners)
Yeah, that's just what infrastructure work is. Like I said, take that analogy, multiply the complexity by 4 (at least... really maybe multiply it by an order of magnitude).
Let me put it in perspective. I've been coding since I was 12, I taught myself C to build a MUD in middle school and high school. I had about a decade of full stack professional experience in Java, PHP, javascript and I'd done infrastructure work with EC2 and chef before. When I moved into DevOps it was overwhelming.
I've been in DevOps for 4 years. I built that equivalent DO infrastructure with Kubernetes just last week (and in a week). I started the week going "Fuck, I don't know what I'm doing." The first 3 days were just spent reading documentation. Day four was spent writing the terraform and kubernetes manifests - with a distinct feeling that none of this was going to work because I was missing several key pieces. Day 5 was spent putting a few of those pieces in place and debugging. I finally got it working late Friday night. I took on a ton of tech debt and made a bunch of compromises just to get something working. I'm not the least bit happy with what I have working and intend to totally rebuild it on AWS when it comes time to build production.
And that's with 4 solid years of doing infrastructure work full time under my belt. For someone with no infrastructure experience? I would estimate 1 - 3 months. There's just way too much to learn to think you could do it quickly and simply.
With an express backend, if you have javascript experience, you really don't have much to learn. You need to learn how http interacts with the backend, how the backend interacts with the database, and databases (SQL). That's it. Learning database is not nothing, there's a lot that comes with it, but that's still only 2 new tools really.
With infrastructure, you need to learn networking, databases, securty, container orchestration (how does high availability work? Scaling?), bash, linux, provisioning, terraform, Docker, Kubernetes manifests, monitoring, secrets handling, and more. And for a lot of these things, the solutions are far from simple or perfect. Even when done as well as can be with modern tech it feels shakey and cobbled together at the end. You're tying a dozen different tool types together to solve a dozen different problems and you have dozens of choices for each tool type.
Like I said, infrastructure is just like that. And it's important to have the right expectations going in to it.
If you can't tell, I've had this conversation with my peers who stayed in full stack a lot.
Meanwhile, going fly.io sounds sensible to me.
At my previous employer (~50 FTE of devs, 2-ish FTE dedicated to infra) Kubernetes worked perfectly fine, and I think in that context it made a lot more sense.
With this approach to hosting and deployment, I think Kubernetes' main advantage is that it opens the door to new kinds of infrastructure businesses, not that it makes hosting a website any easier.
I've tried many of the serverless platforms and maybe it's the types of applications I work on, but I've found most of their limitations (short runtime, limited access to resources on your private network) basically make them useless. The more self-hosted types that don't have these limitations lose out on many of the benefits or are leaky abstractions on k8s.
Cloud Run has all the benefits I want: extremely easy deployment and scaling, as well as the ability to scale to zero if you need it (though generally you don't), while still being able to run basically whatever workload I want. My current employer is mostly a Python shop but we recently deployed a little .NET core service on Cloud Run and it's been awesome.
Source: I'm the Cloud Run PM and we have commmunicated about that publicly in the past.
Do you have any Google docs or blog posts that talk about this?
I always wondered why you need a Serverless VPC connector for "vanilla" Cloud Run (or you have to use Cloud Run on GKE) to access VPC resources, but I suppose this answers that question.
I am slowly moving towards using Hashicorp's Nomad running on Fedora CoreOS using the Podman and QEMU drivers. I rolled out a Nomad at work for internal projects and it let's me get things done quickly without living in a total YAML hellscape.
1: https://docs.fedoraproject.org/en-US/fedora-coreos/getting-s...
3 x Consul server
3 x Nomad server
2/3 x Vault server
It's long since I operated k8s but IIRC I think you can get similar capabilities and redundancy with 3-5 machines?That's before you start looking at actual runner nodes, load balancers, proxies, logging and monitoring infra, etc...
Unless you cheat (which I think many do) or you're big enough, that overhead can be meaningful.
FWIW we recognized this was too much overhead for many users. Nomad 1.3 supports service discovery so you can start without Consul, and 1.4 will support secure variables to get folks farther along without requiring Vault.
So 3 Nomad servers should give you a pretty featureful and highly available cluster these days.
I really don't get why people love that company so much.
We are a small team of 5 infrastructure engineers and previously managed 200+ libvirt VMs running on bare-metal HA hypervisors in a GlusterFS storage pool (software agency, different customer application services). We started to migrate to GKE in 2017 and finished within a year or so.
I know many associate k8s with a yaml mess, but this is actually our most favourite part of it. We are able to describe a whole customer project in this format and it's not something we have to maintain in-house (Ansible). As long as you don't try to be smart (templating/helm, operator dependance), it works out pretty well, prefer plain manifests and extend that with you own validation scripts.
Nevertheless, if you have no 24/7 operations, stay the hell away from bare-metal - go managed.
My company has clients who usually have very simple requirements. A Python/Django app server and a database. Sometimes there will be another background service or two (memcached or equivalent etc).
The most complex site we had was the above but with some Postgres replication clients.
We use docker and docker-compose. We've used ansible in the past as well as fabric and other simple solutions.
We've had a couple of devs try and convince us that we should be using Kubernetes and I counter with "it's overkill for what we need". Am I wrong?
You can install all of those servers in different containers, and then combine them in the same pod. For all intents and purposes from the outside, it will be a singular VM. But from the inside, you will be able to separate all those servers/tasks to separate containers, running from inside the same virtual localhost machine. They can also use the same PVC, making running a stateful app much more easier. You dont even need to make it a stateful set.
You get a lot of benefits with this - you will be able to easily manage each different server in the containers inside the pod. Easily manage their resource constraints. Security. You can make the pod's containers not accept connection from anything that does not belong to the particular app that they belong to. K8 will manage resources in the cluster, its autoscaling up and down, everything. All of the stuff that you had to maintain scripts or ansible to make happen in non k8 setups will be automated.
K8 is basically abstraction of the non-business stuff a lot of infra approaches were doing. Its containers inside VMs without you needing to manage VMs.
You're not necessarily wrong, as long as Docker Compose isn't incompatible with what you're trying to do - e.g. if you'd need overlay networking across multiple nodes, or scheduling things across them in one go, then Docker Compose might not be the best fit and you might instead be better served by looking in the direction of Nomad or even Docker Swarm, though the future there is unclear - maintenance mode project, but very similar to Docker Compose and comes out of the box with any Docker install.
Either way, Kubernetes might indeed be overkill for simple setups, unless you're using just a subset of its functionality and are running lightweight clusters, like K3s or K0s. I guess some might be pushing it because it's basically become the industry standard, at least in some capacity, in some places, or maybe people just want to put it on their CVs.
The reason we went for that setup is that it helped us cut cloud/hw costs significantly (at the start we pretty much had two workers and that was because we ran everything with replicas=2) - each individual site had small requirements, and with k8s we could guarantee enough resources while binpacking as many of them per server as possible.
The actual deployment story can possibly get simpler than docker-compose, but I'd say the real question is whether you'd get a financial win out of it, as it seems you have a pretty good steady state going.
I would be hosting on vanilla cloud with or without Kubernetes.
For my new projects nowadays, I'm pushing mainly serverless approaches using AWS Lambdas (behind API Gateways for stuff that needs to be reachable by HTTP).
I think this shifts the complexity from managing Kubernetes and its accompanying ten-thousand-yaml-files to infrastructure-as-code and the complexities of dealing with AWS. And I happen to prefer the latter, even though it's not infinitely better by any margin.
For the few things that needs to be always-online, or 3rd party self-hosted apps, I'm still on Kubernetes, or pure Docker if possible.
Also how is the cost of Lambda? I know for AWS the logic apps have a really high cost (but function apps seem to be reasonable)
Not to mention any managed services you want to use, which also will lock you in. So I don't do anything extreme to avoid vendor lock-in, other than making my code general enough to only have a small surface area for the lambda entry point. As a practical example, all of my APIs hosted on Lambdas are ordinary ASP.Net apps that would work identically if hosted in Docker containers.
Pricing so far is one of the biggest benefits of doing serverless approaches. I'm down to paying a couple of dollars per month for something that I'd pay tenfold for if doing Kubernetes. Both the monetary sum of only paying for what you're using, and also not having to worry about cluster maintenance, scaling and management is a godsend.
So cloudformation YAML? or CDK?
I went from on-metal K8s clusters, which were a complete PITA and required a full team to manage, to using EKS which has been everything K8s should be... easy peasy.
We don't really have a use-case for Boundary but it looks pretty neat as well if you do.
Was on k8s for years and I don't miss it one bit.
While there definitely is some complexity once you get serious and set everything up properly with raft, federation, Connect, CAs, proxies, ACLs, proper secrets lifecycles... I find it's worth it. With the current assumptions that HC will keep improving and existing bugs and edge-cases will be ironed out.
Also: our main reason to adopt Kubernetes was to stay cloud-agnostic, but we soon realized that this is as unrealistic as writing a complex app's SQL in a vendor-independent way.
Instead, we decided to embrace our cloud (AWS) by using their CDK tooling and leveraging their features as much as possible. If we ever need to switch to another cloud we will bear the cost then, but for now it is clearly YAGNI.
Same goes for heroku/digital ocean app services. Even elastic beanstalk. If you are large enough that you need to manage your own k8s cluster, that is one thing, but I would encourage you to look at your needs from a usage and compute perspective long before you start solutionizing with trendy technologies.
I am trying to move on k3s but it is just too complex to run anything and there is still not solved problem of exposing services to internet.
What I want is to declare I want this service to be under this domain and this IP - so for that you still need to configure your load balancer (bare metal) manually, setup certificates etc. I am writing a tool to automate this, but it's been a pain.
After initial setup you can do it quite easily.
Exposing a service on selected domain is several lines in Ingress and adding certificates is several more. Example: https://cert-manager.io/docs/tutorials/acme/nginx-ingress/#s...
What I have in mind is an external server that is not being a part of the cluster that bears the role of load balancer. It will contact the cluster and look for services and then setup up a reverse proxy based on their declared hostname, then setup certificates and update DNS records at DNS provider.
As far as I know something like this does not exist.
Maybe Traefik has such a capability, but their documentation is so complex I have no idea.
Yes, I need to add A records with IPs for each domain, but that's one time setup. I did it manually, but you can automate it [1] (depends on what you use for DNS provider but you can extend it to support your provider or maybe there is another existing solution).
I'm not sure that one server in front of the cluster is more reliable than using all cluster nodes for load balancing. I guess that in automated solutions like [1] cluster's node could be automatically deleted from DNS if it went down.
My setup is not so big so I don't have real need for load balancing, but it seems possible with existing solutions.
As for DNS records, external-dns[2] works perfectly as long as your DNS as some way to doing automatic updates.
1. Is it really so complicated?
2. Is that complexity incidental or essential?
3. Could we get away with a simpler set of abstractions for 90% of applications?
1. You read the O'Reilly book first (or another good book). There are a few unexpected abstractions (replicas, services, deployments, etc) which the book explains nicely.
2. You pay for a hosted Kubernetes. Google's is great. EKS is workable, but you may need to spend more time configuring it.
3. You don't mess with the networking system, and nothing goes horribly wrong.
Our clusters peak out at close to 400 CPUs, and Kubernetes generally does what it says it will do.
One caveat: If your app can be deployed using a "platform as a service" (Heroku, Render, etc), that's usually a better idea than Kubernetes. Kubernetes makes sense when a PaaS starts feeling too limited.
Most questions seem to revolve around a tiny part of the puzzle, or a small "just starting out" phase and completely forgets about the lifecycle of the business process that it is built for, and the existing systems it needs to interact with. Even a startup will have that problem considering most are trying to get bought which essentially means being absorbed into a legacy company. So even starting out with no legacy to worry about is just a stay of execution.
2. Essential. K8s solves a problem that's quite complex. You can't really solve it in a simple manner.
3. Probably, but that 10% will require the additional abstractions and complications anyway, and it'll be easier to manage one system rather than 2.
1. It's only as complicated as you make it. Kubernetes is essentially PKI (which is a must in any case), a REST API, and a scheduler. It stores some stuff somewhere, and you can add more stuff for it do have more features. I wouldn't call that complicated and it's essentially what Swarm and Mesos do as well (minus the PKI part).
2. PKI is essential. If you think that's complicated that's a whole different problem. Everything else is incidental. If adding more OpenAPIV3 schemas or REST API seem complex, again, not really a Kubernetes thing, mostly a general software development thing.
3. Yes, as 90% of applications really only exist as mediocre CRUD viewers you could run on a potato. Also, 90% of applications don't need to be as highly available or scalable as people might think. Then again, ecosystem complexity in software development combined with the lack of general knowledge (i.e. how to use an RDBMS properly) makes that while the software is simple and could be run as a single statically compiled binary, it generally is a mess, requiring more messes to make it run. But since that is cheaper (less developer time spent, more cheaper developers available to do that type of work), that is where we end up.
2- See 1.
3 - If you can containerize your app in a simple way, then yes.
Note that a stateful app that would require attention in a bare metal server or a singular VM would still require that kind of attention on K8 as well. K8 just removes the need to manage the VM infra. And makes running your infra as code much easier.
If you need to run a stateful app in a highly available manner, you can do it in K8 and it would be good - however you will spend a similar effort for maintaining the high available services like you do in other venues. Ie, if your stateful app requires a Percona cluster and a NFS cluster off of K8, you will still need to launch and maintain those services. K8 operators make these a lot easier to launch and maintain. But its still maintenance nonetheless.
Using managed, hosted databases can work for the database part. But they are expensive. So launching a database cluster via a K8 operator would be cheaper to maintain. NFS is a problematic thing across all platforms. So if you need it, you either launch a rook-ceph cluster to provide a shared filesystem or use a hosted service like Google File Store.
Even hosting redis etc.. is really straight forward.
It is funny, but the complexity starts to happen where you want kubernetes to handle other stuff: like hosting databases, or other storage resources, and if you want to for some reason I will never understand have your external services essentially communicate directly with kubernetes rather than have some middleware service you pay for do that for you (like a load balancer, etc.)
One thing I did have an issue with was setting up SSL... that was surprisingly stupid. Should have been much easier to do that with LetsEncrypt.
Then you run into a litany of issues with networking (like you mentioned SSL termination) and stateful apps or databases.
Even in this thread, someone mentioned how Redis defaults lead to a lot of issues in containers.
But then, someone is trying to fetch 50GB files and now you need to play with buffers. The script misbehaves so the API rate limits you and now you need to handle credentials, back-off, etc. The script hangs in some strange state and you need to add structured logs to figure out what is happening. Now we need to upload multiple files in parallel, are we going multi-process or multi-thread? Is python the right language? Are we going to use one pod or many?
See how it quickly gets complicated? Add to all that the fact that it’s easy to spin-up rabbitMQ with some defaults with helm locally. So you do that in production as well and when it goes down you don’t know what’s happening.
As another commentator said, there is a level of knowledge that is required to things in production reliably.
... and we're now fully back to the mainframe era with the people in white coats who "run" the computer.
The cloud truly is mainframe 2.0.
Depends. Are you a >500 Developer org with many services? Then it's easy compared to what's out there. Anything less than that I'd say it's complex and you'd be better off using a PaaS
> 2. Is that complexity incidental or essential?
Depends. If you're going to do simple things forever then it's an overkill. But if you expect to grow in unknown ways in the future and don't want to waste your time doing bunch of migrations in the future then it's essential.
> 3. Could we get away with a simpler set of abstractions for 90% of applications?
Maybe? Heroku, AppEngine, CloudFoundry tried, but didn't go too far. Let's see what new crop of PaaS offerings are able to do
Kubernetes is available as a managed service in AWS, Azure, Google etc and this is likely to be the most popular deployment model.
By any definition this is a PaaS and if you add in custom monitoring, logging, security, ingress etc. is going to be just as simple and significantly cheaper than using a managed solution.
If you're just building a basic website then sure it's an overkill but fewer people are building those these days.
> 3. Could we get away with a simpler set of abstractions for 90% of applications?
Yeah but you can do that in k8s too, check out knative serving for example. K8s encourages the creation of higher level abstractions, with the advantage that you always have the break glass to dig into the primitives, which you don’t get with a lot of other systems.
2) necessary at scale, incidental before then.
3) yes.
b) Kubernetes clusters can span multiple accounts, clouds etc.
But it doesn't get brought up as an option very often because docker basically FUDed themselves by having two things called swarm and then loudly killing the older one making everyone think it no longer exists.
It's not for hyperscalers and it's got a limited feature set compared to k8s, but it's simple enough that you can really learn how it works and how to make it do what you want even if it's only a small part of your job.
If you just want redundant services, zero-downtime upgrades, and either manual-only scaling or very restricted autoscaling, Swarm is likely sufficient.
Main downside? Not available as a managed service, at least from major providers. Then again, if you're OK with managed services, you would probably prefer either a fully-managed PaaS (Heroku, Azure Web Apps, etc.) or a managed k8s.
Didn't pass my BS test.
I am glad that people are moving on to something that will exhaust their creative juices on something ... pointless instead of focusing on delivering value for their customers :)
More people using the brand new tech - less competition :)
To be honest, even with the technical overhead it'll probably solve a lot of problems for us from a workflow perspective. We've (the engineers) been arguing for more component-level testing for years (as opposed to the all-up E2E testing we're required do now, which typically turns into component-level testing anyway), and containerizing everything is a good excuse to push it into reality. It'll also make deployments a lot easier (just roll back to X image if there's a problem). Right now we have tens of thousands of lines of hand-written deployment scripts that manage everything and have to be maintained, and intimate knowledge of how they work is often limited to who wrote it (many of whom are no longer with the company), and if there's a problem you have to do surgery on the environment. Kubernetes will give us a unified deployment architecture with problems you can google.
Also from my experience, people will start complaining as soon as the new deployment with k8 start failing and they have to fix it.
But its a good opportunity to make the transition to more stable architecture.
My suggestion is to take it slow and do changes one system at a time. Start with stateless application with less risky deployment and as you learn move others.
We have about 100 devs in multiple teams. Kubernetes provides great level of standardization and transparency - completely different experience than VMs, where admin team had too much ability to cut corners and build technology debt. People would riot if they had to go back to these days.
A few warnings: * It takes some resources. Maybe can be mitigated with k3s or similar, but I don't have first hand knowledge here. * It requires some time to learn and configure properly. If your entire team is 3 people and you are on limited budget, probably not a good idea. * Adopt some tools (helm?), standardize deployments, where possible. Bare k8s is bit too much for daily work. * Read good practices and don't try to be smarter, at least until you really know what you are doing. Limit misconfiguration may really burn you at least convenient moment.
And then a couple weeks ago, I was tasked with standing up a new Ansible AWX server, which now done via a k8s operator. It was an exquisitely painful experience. This is potentially a bad example because I'm pretty sure IBM's plan with AWX is now to make me suffer, but through that entire process, k8s just felt like extreme overkill.
I'm pretty sure that's going to be the last time I use k8s. I know it makes sense for some use cases, but I it just doesn't feel intuitive in any way. And although it may seem more efficient, I absolutely dread having to troubleshoot any problems down the road.
I'm probably not the target audience, but thought I'd leave a comment for fun anyway.
Do I recommend Kubernetes to other people/companies though? Absolutely not! The learning curve is incredibly steep, and it really does take investment into understanding how it works.
But to anyone who is looking to use Kubernetes, I highly recommend https://helm.sh since it actually makes templating deployments significantly easier.
But in reality, I think I developed an allergic reaction to complexity and hype. I took some metrics; things like recording the time taken, steps taken and happiness generated from my current build/release stages, then comparing to k8s.
In conclusion, struggling to learn k8s forced me to find joy in the simplicity - knowing that one day (that will never come), I can just hire someone to do this... "It's only a problem when it's a problem".
For now, I have a lovely bash script that is triggered on Github releases (using Actions), which uses doctl to do the following:
1) Create a new server from my baseline image 2) Run the setup steps as defined in the Dockerfile, although it doesn't use docker (it just makes sense to keep the configuration I used to have) 3) Copy the built-and-tested version of the repository to the new server 4) Run any post deployment scripts, like database migrations, whatever 5) Move the reserved IP to the new server
It takes about a minute from me clicking "new release" in Github to seeing the changes hit production. If there's a problem, I move the reserved IP back. Load balancers, database clusters, etc... they're all set up manually because "it's only a problem when it's a problem".
Kuberneeties only ever generated problems for me.
We feel the same about Docker. People have no idea what's running when they download a docker image. Stuff can be buried deep deep within an operating system image. Security should be simple, transparent, and minimal so it can be reviewed easily. Reviewing a docker image is impossible. I'm convinced the correct place for isolation is systemd. This guy wrote a great starter for hardening the crap out of your services: https://docs.arbitrary.ch/security/systemd.html Systemd offers a bridge too with nspawn if you're not ready to undertake ultra minimal hardening of services.
Scaling is a "sexy" problem to have though, and software "engineers" love to think that their SAAS product with 100 users is going to take Google scale workloads; thusly what could be done in a LAMP stack on a single DO server, is inflated into a fantasy that will never come to fruition.
This is true of software packages and especially of third-party libraries; supply chain attacks are supply chain attacks. But similarly, supply chain controls are supply chain controls, and using Docker does not mean running someone else's container.
(For example, we build our own hardened base images, and on those we install our own services, and the result is precisely as trusted as building our own hardened AMI and installing our services on that.)
I have spent the past couple weeks working with kustomize since I do not like helm and while it gets the job done I think Tanka would be better.
We are on GKE which makes things a lot easier and I personally would not choose to run my own cluster.
(disclaimer: i worked at hashi for 4 months in 2020 but not related to nomad)
I still run KEDA at home for managing plex, home assistant, some game servers and other of my own projects. But being the only one who is using the cluster is a different use case than getting RBAC, ingress and management set up correctly for a production cluster IMO. I’ve never had the sole responsibility or permission over a cluster before, so it was a daunting step I decided not to take for my own sake
People talk about it being incredibly complex, and honestly I don't see it. Yeah there's a layer of jargon you have to dive into, but it all makes sense once you start building something with it. By far the most complex pieces for us are the integration points with AWS (we're using EKS.) The examples/docs available are just not that great.
We run most of our app on Google App Engine explicitly to avoid devops work. However, we have a stateless-but-memory-hungry image manipulation service that was just too expensive on GAE. We migrated that service to k8s on Digital Ocean.
It was a disaster. I mean, it worked, but suddenly we were spending a lot of time learning k8s and fussing with k8s and it slowed down feature development. K8s is a time sink. So we migrated the service to Digital Ocean App Platform and velocity returned to normal.
I'm not wholly thrilled with DO App Platform. It has some maturity issues, and while it's cheaper than GAE, RAM is still more expensive than Elastic Beanstalk (which charges you more or less the EC2 VM cost). So we'll probably move it there someday.
i would stick no matter the company size on IAAS + $Deploymenttool (ansible or so) and docker and then get comfortable with and only then, when everything works as intended make the switch to k8s.
creating a scalable system is complicated within aws account limits
all we really want is to shove docker containers behind a load balancer and not worry about having to manage yet another system
At work, we're currently trying to migrate our stack to k8s. Why ? Because our startup is getting bigger and bigger and our current platform sucks, but our products are becoming a lot more complex as time goes. We benched a few platforms and landed on EKS + ArgoCD + Vault. Works really well.
For personal projects I roll with just docker.
I'm bothered by the minimal requirements of k8s, I want to deploy on 5$ machines
I used it to host several small projects in cheap virtual machines. The setup is very straightforward. I guess we just need better editing support of YAMLs.
Thanks for the recommendation!
or just containers on the virtual machine?
I would love to deploy with docker and no other orchestration tools.
We came from Ansible managed deployments of vanilla docker with nginx as single node ingress with another load balancer on top of that.
Worked fine, but HA for containers that are only allowed to exist once in the stack was one thing that caused us headaches.
Then, we had a workshop for Rancher RKE. Looked promising at the start, but operating it became a headache as we didn't have enough people in the project team to maintain it. Certificates expiring was an issue and the fact that you actually kinda had to baby-sit the cluster was a turn off.
We killed the switch to kubernetes and moved back to Ansible + nginX + docker.
In the meantime we were toying around with Docker Swarm for smaller scale deployments and inhouse infrastructure. We didn't find anything to not like and are currently moving into that direction.
How we do things in Swarm:
1. Monitoring using an updated Swarmprom stack (https://github.com/neuroforgede/swarmsible/tree/master/envir...)
2. Graphical Insights into the Cluster / Debugging -> Portainer
3. Ingress: Treafik together with tecnativa/docker-socket-proxy so that traefik does not have to run on the managers
4. Container Autoscaling: did not need it yet for our internal installations as well as our customer deployments on bare metal, but we would go for a solution based on prometheus metrics, similar to https://github.com/UnclePhil/ascaler
5. Hardware Autoscaling: We would build a custom script for this based on prometheus that automatically orders servers of Hetzner using their hcloud-cli
6. Volumes: Hetzner Cloud Plugin, see https://github.com/costela/docker-volume-hetzner - Looking forward to CSI support though.
7. Load Balancer + SSL: in front of the Swarm using our Cloud Provider
Reasons that we would dabble in k8s again:
1. A lot of projects are k8s only (see OpenFaaS for example)
2. Finer grained control for User permissions
3. Service Mesh to introduce service accounts without requiring to go through a custom proxy
Honestly - I even use it personally for my self-hosted stuff at this point. The learning curve is... steep. But once you come out the other side, it's a great tool.
- Ansible to provision such VMs
- Docker to start/stop containers on such VMs
It feels like a breeze of fresh air!
AWS has EKS for a reason.
Now building fully with serverless
Therefore you either can’t find anyone or more likely you hire less good DevOps engineers.
The solution is to not use k8s as a startup. The less a DevOps engineer can shoot themselves in the foot the better.
it all scales and works fine. there is one or two problems around but not enough for me to consider it does not work.
Would I use k8s just for static websites or single API? No. Would I use k8s, for rarely updated solution, where low costs are #1 priority? No. Would I use k8s for complex microservice architecture with a long list of ever-growing implicit/explicit requirements and a lot of moving parts? Definitely, because now you just need to either use some built-in k8s feature and/or use CNCF eco-system to supply you with almost anything you need.
Kubernetes gives standardization, it's good in enterprise, where high complexity and poor communication is normal. It covers a lot of typical requirements for applications and you can learn a lot about solution just by looking at k8s cluster. However it's a time sink, if you really want to learn more about k8s and CNCF eco-system.
This is a major problem if the team isn't well versed with devops tools (which is often the case at smaller companies) and can lead to lots of issues pushing a lot of work to a devops resource (or team) which requires setting up a separate dev/staging cluster.
I think the preferred alternative for companies that struggle with this overhead is to use a slightly more expensive managed service, especially if you're just developing a typical MVC/MV* app.
The two philosophies at megacorp here seem to be "I built it from the ground up to target X service" where X service is usually amazon serverless or something, and "I built it in docker containers but I don't know about the cloud".
The former is a conscious decision, and we (Architecture) have a serious, sit down discussion with them about what it actually means to be fully cloud native for that particular service. This discussion ranges from cost analysis, to things like "is your application actually build correctly to do this", to "you're not going to have access to onprem resources if you do this", even asking them simply "why".
A lot of the time when the teams realize they're going to be on the hook for the cost alone they back out, and a lot of teams try to do it because "we don't understand K8s". Well, it doesn't get much better in Cloud Run either folks because you're trading K8s yaml for terraform or cloudformation.
Where it has been successful is for teams which own APIs which only get called once a month, or very low traffic APIs. I hate to say it boils down to cost, but a lot of the time it really does boil down to cost.
Additionally we've seen a weird boomerang effect as clouds offer K8s clusters which are simply priced per pod rather than per worker node (like GKE Autopilot). A lot of teams which straddled the middle of "low traffic but not low enough to really migrate" have found they're quite happy in GKE Autopilot. They use autoscalers to provide surge protection, but they just use Autopilot with 1 or 2 pods running and it keeps the costs down. That also means we can migrate them to beefier clusters in a heartbeat if they get the Hug of Death or something from HN. ;)
The second use case I discussed gets railroaded into our K8s clusters we built ourselves because we can typically get them to use our templates which provide ingresses and service meshes and the developers don't have to think about it too much, and the Devops team is comfortable with the technologies. While it means that there's a bit of "rubber stamping" and potential waste, it's allowed us to use K8s and the nice features it provides without having to invest too much in thinking about it for an individual application.
Personal; I use docker-compose on VMs
For context, I ran a DevOps team for the last 4 years that managed two products on AWS - one on EKS and one on ECS. I was mostly managing by the time we had k8s in our stack, so I didn't get to interact with it much directly, but I know infrastructure generally (and I know ECS inside and out, unfortunately). For that infrastructure, we had a whole team managing the ECS deployment. We managed the EKS infrastructure with the equivalent of one DevOp's time or less for years. It was only when it started scaling to millions of users that we needed to give it more time and attention. Both infrastructures (ECS and EKS) were pretty complex with multiple services that needed their own configuration and handling.
I left that company a few months back to try to build my own thing and I just finished building out the alpha infrastructure for it on Kubernetes. I can now safely say, as an infrastructure engineer, Kubernetes is an absolute joy to work with compared to lesser abstractions. At least, when someone else is managing the control plane for you. It has exactly the right abstractions, with the right defaults, and it behaves basically exactly as it should.
Yes, it's complicated. Yes, there are a lot of moving pieces. Yes, there are hard problems. That's just the reality of software infrastructure. That's not kubernetes, those are just the problems of infrastructure. There's a whole set of problems kubernetes is working to solve in addition to those. Remove kubernetes and you still have those problems, but then you also have the whole set of problems Kubernetes solves as well.
I think what's really happening with this whole "Kubernetes is too complicated thing" is that a lot of teams expect to be able to use it like Heroku. That's not what it is. Or they try to build out infrastructures with javascript/php/python/etc engineers. You wouldn't try to have a team of front end engineers build your rest backend. It's not reasonable to expect javascript engineers to know how to build and operate an infrastructure - at least not with out dedicating themselves to learning the tooling and space full time for a while. Think of it from the perspective of a frontend engineer learning Python and Django to build out a rest backend, and then multiply the complexity by 4. That's just infrastructure regardless of what you're using.
If you just need to run and scale a container fast and simple, with maybe a single database - then sure the PaaS providers might fit your needs for a while. But eventually the trade off is going to be cost and limitations. You'll eventually need a piece of infrastructure they don't provide.
TL;DR Kubernetes isn't the problem here. Infrastructure work is just plain complicated. If you want multiple services, high availability, reliability, scalability, security, and performance, it's just complicated and hard. Don't short change it. Dedicate someone to learning it or hire someone who knows it.
It's like asking people if they have moved from VLANs to duplicated flat networks because that is "better". You're just trading one complexity for another.