I don't have services like Pingdom or Pager Duty for my sites. I have Rackspace. If a site goes down, they immediately attempt to recover it, acting within the account rules I set up with them. If that fails, I get a phone call from an expert and we troubleshoot together.
I once got a call that one of my servers had started sending out a high volume of email. When I took the call, Rackspace folks had already found the nasty script doing it and just needed my ok to take the site down for a few minutes to boot it.
The service is really great. But, it ends just below the application layer, and these days, that's where most of the problems show up. So I am looking at moving off of Rackspace, but I'm looking up the service ladder at application-aware hosting services like WP Engine or Pantheon. As opposed to down the service ladder, to a self-service "don't call us" system like AWS or Google.
Their support used to be the best in the business and we were willing to pay a premium for that but even their support started to suffer starting early 2015.
It is unfortunate because Rackspace was an amazing partner to our business for a long-time, which made it a very hard choice when we decided to move on.
We were very expensive but the companies must have saved money in hiring their own Engineers as were always growing.
AWS/GCP/Azure have made doing a lot of what we used to do incredibly more accessible to more people.
90% seems way off. AWS is known for not being cheap (neither is Rackspace), but I highly doubt that Rackspace was ten times the price of Amazon.
There are MANY applications that simply are too complex to handle such an arch... so, - do so, when you can, but know when that's just stupid.... stacks are like snowflakes, we LOVE them to death, or we love them to DEATH...
Not if you need always-on instances.
customer support
Maybe there was a way to escalate it, but it wasn't readily evident. Which tells you something. So, I just moved on to AWS. More transparency on the platform and more predictable support.
Bad news. AWS is no different. EC2, IAM, whatever will have issues for hours, AWS won't update their status board (to the point someone wrote an extension to show the "real" status [1]), etc. And this is with business support we pay for. You needed to autoscale? Better go make yourself a mojito and wait for the dust to settle.
"There is no cloud, its just someone else's computer" [2]
[1] https://chrome.google.com/webstore/detail/real-aws-status/ka...
In the end someone has a pager, a machine and an editor. In ye-old-days this problem was in-house with a COLO machine. That's the reasoning behind uptime. 99%/year is 361 and just over half a day. Can you loose two and a half days at christmas? (web commerce) or the same time on a new product release? (SaaS) or 2.5 days development time at a critical bug fix with your customers?
The bad stuff always happens at the worst time. The market is saying, ^cloud^ but the tradeoff is ^service^. I don't know the answer(s).
but such a product would be expensive