I think we finally realized that fixing something in prod via ssh is not a good solution and might introduce new bugs on its own. Rather build an infrastructure that allows you to rollback fast. It is also not worth to fix individual machines in a big cluster, just throw them away and bootstrap them from zero. This way you make sure that you don't have accumulated patches and workarounds on your nodes that might lead to future failures. In some companies we reached the point where you don't even fix a cluster in a multi-cluster setup, but throw away the whole damn thing and bootstrap it from zero.
Agreed. Reminds me of articles and blogs that GitLab team publishes these days.
Granted it's not going to be easy to have a platform where every tool can seamlessly connect, but we can start with baby steps such as services that automate multi-cloud deployments like the article suggests.