It's been fascinating watching how dramatically that viewpoint has shifted over the years to the point where it is now a novel idea to do so.
It's been fascinating watching how dramatically that viewpoint has shifted over the years to the point where it is now a novel idea to do so.
The level of convenience that cloud providers give is just orders of magnitude more efficient and easier.
On the flip side, I wonder how much better the support from a cloud provider is if it's an isolated problem and not something that's setting twitter aflame.
If it takes a cloud provider in the order of hours to get me back online, I could probably get the same sort of service from one of the better colos/hosting providers, especially if they were local and I had the ability to make a call to get support.
There are other conveniences to cloud providers of course, but I think I if I could find highly skilled ops people and pay them well, I would run my own servers every time. For the kind of games I've worked on, the money/CPU cost of cloud is ludicrous.
The trick these days is even finding high-level ops people who aren't already working 3-400k jobs for AWS/Azure/GCP
I’ve had 4w to resolution for major service degradation on a big name provider and ~2d for full site outage (highest support tier)
You said it yourself - aws, goog, msft have very good sres but they’re not your sres. Meaning they dont care if you have a big event/demo/deal close coming up before they start doing network gear upgrade and such...
The other consideration is that things generally don't just 'go down' like they do with bare metal (because HA - replication and so on), but if they do, it's likely affecting a large portion of the internet too.
I don't even know if these salaries are truthful now.
Those are certainly in range for DevOps people at the higher end of the scale.
In the same way most people/companies don't really need to care what hardware their application runs on, only that it meets some bar of quality/cost that's appropriate for them. If someone else is delivering this then you've removed a small department's worth of overhead/planning from your corporate structure.
So, not a ideal analogy...
The sibling comment pointed out that vendor lock in can be a problem which I agree with, but I think for most of the industry that's a problem of protecting yourself from predatory price hikes/services being deprecated rather than the problem of actively pissing off people you need.
I've had zero downtime due to running on a Digital Ocean VPS the past few years and very, very brief outages due to my own decisions around various upgrade and backup decisions. I spend less, I get full control and there's zero risk of a surprise bill.
I think the key thing is that most customers won't blame you if your service goes down due to an AWS outage.