It's not like the cloud is a magical place where no unexpected issues ever happen. Cloud providers can be surprisingly buggy, especially AWS, particularly at scale when you start hitting the implicit scaling limits their docs conveniently forgot to mention.
Each technology has their unique challenges.
An extra abstraction layer like k8s makes it a lot easier, which is exciting. It's also precisely why IBM bought Red Hat - a well-built k8s distribution like OpenShift is one of the very few real alternative to public clouds for many companies.
OpenShift is built on a prescriptive mentality. They choose everything for you from OS through pipeline and many deviations are unsupported. As you mentioned there are others, but depending on your environment OpenShift is either a very good fit or square-peg-round-hole.
It's not like I'm personally afraid of k8s or building docker images myself either, but anything to help my team be more productive without needing to hold hands or do gruntwork on their behalf improves the lives of everyone.
Running custom k8s in production is the 2019 equivalent of using Gentoo.
0: https://www.joelonsoftware.com/2002/11/11/the-law-of-leaky-a...
Yup. We've run into that repeatedly. The "we didn't notice anything on our side, please send more screenshots and logs" gets really old when working with "managed services".
Network packet losses/truncations, EKS control plane failures, cloudformation stacks getting stuck in really wierd states, inconsistent cloudformation implementations for new and existing services, ENI weirdness in containers...
Managed services just feel like they aren't - more and more every day. Amazon (or any other cloud) will never have the same investment in your availability and infrastructure as you will.
For a team with the goals of IAC, hands-off, and high-availability infrastructure, migrating is a pain.
It would definitely be an annoying migration, even if import works for your use-case, but honestly, dealing with CloudFormation is so irritating on a day-to-day basis for me that I'd consider it worth it.
The problem is that the reality has not lived up to the promise, and "first party support" means "only the first party can support".
After a few rounds of back-and-forth, I eventually gave up.
(I even gave them permission to copy test snapshot data, at their request, and sample queries to reproduce the issue. They could have done all the testing they wanted in-house. But that might have actually required them to spend time on their side to diagnose and fix something, I guess.)
It wasn't a critical issue, exactly, and we worked around it in code. I mostly just wanted to report a problem.
Maybe you simply haven't seen this kind of operation in action before?
This is all operating on the assumption that those entities will care enough about your problem enough to do something meaningful to fix it. As others have stated here, at least in the context of "Big cloud providers", that frequently is not the case. Often you just get the runaround (continuous delays, requests for "more information", other attempts to stall in first level support).
Add a third party "managing" some piece of your company's infrastructure via a cloud provider into the mix and it often gets even worse.
It is true that there are real costs to having your own hardware/software onsite and people who know how to manage it. However, the promised reductions in cost/hassle of moving things offsite are frequently offset or even exceeded by the costs/hassles you get by not having control of things yourself.
All that assumes they even admit the issue is their hardware not your code, which is a mighty big assumption itself.