How do you support tens of thousands of GPUs in a cluster?
Now, I am building different clusters for AI.
I find myself agreeing with the part about "scaling", meaning a (good) DevOps team makes it easier to integrate new projects into an existing infrastructure that's built in house. The point where relying on the vendors above becomes more expensive than having a full time DevOps team only happens once you react a certain scale.
A separate DevOps team --> corporate rebrand of the separate Ops team --> no real change, or even slower because more hoops to jump through.
DevOps was always about meant to be about breaking down the separate / silo'd teams and combining them together. Ownership and responsibility is with that single team for their code from development through to production.
The corporate rebrand of Ops teams is what you're coming up against. It does my nut in and I'm sorry to that you have to deal with it.
Though “scale” seems to mean 10 concurrent users supported by a Kubernetes cluster of 50 systems in order to host a simple web application…
We didn't really need the scale. We just didn't want to deal with any of those problems. Often there's a little checkbox you can check that solves one or more of these problems, the cost of checking the box is dependent on the size of the data itself. This is convenient _and_ predictable. We like that.