Good point.
Hardware aside - thats largely moot on both Heroku and AWS - the 5x cost of the server will very quickly outstrip the cost of a good developer/operations guy. It seems like a good idea - I initially just throw money at a problem too! - but at some point it does not make financial sense, which is what I was getting at.
Note that we don't have every engineer on staff capable of bringing up a new server/rebuild a damaged mysql replication setup, but we do have engineers that can:
- ssh (or attempt to ssh onto) a dead instance
- tail a log
- check disk space, memory usage and cpu load
- check if something is running
- restart something that just randomly stopped
All that stuff is pretty basic, and will get you 80% of the way there - the other 20% being experience. I guess at some point you are paying for the experience of working with a datastore, so there it makes sense.
As far as early warning etc., that does take time, but it's not the big deal it appears to be.
EDIT: I can't math, and 80 + 10 = 90, not 100. Good thing I'm not a data scientist :)