> What we seem to forget all too easily is that the cloud is, in its simplest form, just someone else’s computer.
This isn't even actually true for Infrastructure as a Sevice (IaaS), in which case there's significant intellectual capital that's been invested in the automated provisioning of resources for customers. This is much more true with SaaS products.
In this case, the customers of #NA14 did experience downtime. But they also got the full expertise of the salesforce engineering organization (which clearly is far from incompetent) being brought to bear on the problem, at no extra charge to them as a customer, as well as having likely already mitigated instances of this problem before. These prior mitigations, of course, are met with much less fanfare than an outage. It's hard to appreciate things you don't experience.
> This isn't even actually true for Infrastructure as a Sevice (IaaS), in which case there's significant intellectual capital that's been invested in the automated provisioning of resources for customers. This is much more true with SaaS products.
Does all the "intellectual capital" (read: aws/etc runs a highly complex system made up of various parts that sometimes all just shit themselves when a single low level thing hiccups) make it somehow not someone else's computer?
The point of "it's just someone else's computer" isn't that it's just a computer like you or I have, it's that it's not some magical, mythical beast.
The greatest achievement of aws/et-al is convincing developers they will need "web scale" infrastructure at the drop of a hat.
To illustrate this point succinctly, there are probably thousands of "somebody else's computers" sitting in a dell warehouse somewhere that are no good to me. The cloud is only useful if there are additional services to help me consume that computer.
I'm not saying it doesn't run on somebody else's computer. I could, though, since private clouds are in fact "clouds" that are not someone else's computer.
Again, what I'm saying is that it is not a sufficiently specific generalization to accurately describe "cloud." The definition of "cloud" is fundamentally more complex than "just somebody else's computer."
http://www.pcworld.com/article/160153/google.html
http://news.softpedia.com/news/Longest-Gmail-Meltdown-in-His...
http://arstechnica.com/information-technology/2012/12/why-gm...
http://www.computerworld.com/article/2486914/web-apps/update...
Keep in mind, when comparing it to your experience, is that not all customers are necessarily down when there's a problem.
I'm not dumping on Gmail, running a giant system like that has got to be extremely challenging to get right.
But let's not wear rose-tinted glasses about the cloud. It can be great, but for some applications you can get better reliability by competent people deploying reliable software on reliable hardware.
This doesn't jive with the tech industry hype train though, hence, it being "myopic" to consider running your own services internally.
Is there a list of outages on other 'major' services of this magnitude?
The general trend I follow is that you're good at things you do regularly. That means that internal systems tend to be much worse at recovering from failure (i.e. the time to recover can be days or even weeks) unless they are careful to design with redundancy in mind and practice it regularly. That's a time-honored tradition in large IT shops but it's also one which often requires constant effort to justify the "wasteful" expense even in places which should know better.
Case in point: It's informative that you refer to "mail servers". The plural makes it clear that the company in question isn't in the same league, size-wise, as the target market for most SaaS vendors.
Typical companies, though, are mostly Windows shops and really are looking to get something more along the lines of an integrated communication/collaboration/scheduling tool like Microsoft Exchange or GMail.
This is exactly it. Even if not intentionally designed, it ends up that way. As andrewstuart2 mentioned above, "[customers] also got the full expertise of the salesforce engineering organization (which clearly is far from incompetent) being brought to bear on the problem, at no extra charge to them as a customer". This is great I guess, but if your home-grown solution doesn't use, say, Oracle, to start out with, you're going to eliminate a heck of a lot of necessary engineering expertise when things go wrong. Break up systems even more and suddenly just one generalist can do it all on the side...
Cloud services just have to be equal to internally run ones in this regard, since they save so much in other areas. So I guess people have to decide: is 20 hours of downtime every few years worth running your own services internally and paying people full time to manage them?
It sounds very much like you're assuming the answer is no.
In my entire 10+ year career I've never had a service I was working on have a 20 hour outage.
So, probably, the answer is yes.
Your homegrown service is unlikely to have a 20 hour outage, by being much simpler. Whereas if you use any cloud solution at all, you become a part of a huge and complex system that is likely serving orders of magnitude more requests than your service. And that can go down by reaching situations that would never be seen in your service.
The flip side is, your service is probably peanuts to them. Should you grow using your own tech, you'll hit the same pain points that they have, in their very beginning, probably even during the MVP. They have already taken care of that so you won't have to.
I've seen cases where a single downed router chopped of a whole building of business people who - since everything was on network drives - couldn't do anything better than play solitaire a whole business day long.
Through luck or careful planning and good resources?
Once you have enough hardware to handle load, it all really comes down to risk mitigation.
Assuming it handles load, your service could probably run on a single server, with a single drive, on a DSL connection for years without any issue. I would say if that happens, it has less to do with your skill as an operator and more to do with the luck of not having a power outage, drive failure or backhoe operator dig up a line.
Can your service withstand your office/datacenter burning down? Or your datacenter being shutdown with all assets seized by the SEC (true story)? Or a massive earthquake or flood or hurricane that disrupts the entire city?
I did have an issue once where a very large power outage that took out a data center. Total downtime was about 50 minutes, which consisted of:
1. Calling the hosting provider and ascertaining the problem.
2. Pulling the latest code onto the fallback server (from a different hosting provider).
3. Restoring the database from the backup server into a newly-created database server.
4. Changing the configs.
2 and 4 were able to happen while 3 was running.
As an aside, this is largely why specialized cloud services like AWS or AppEngine that promote vendor lock-in seem crazy to me. AWS can and does go down, and if your infrastructure is built around their tooling, you can't just provision a new box on a new provider.
How arrogant. I seriously doubt the services you've managed have even served a fraction of the data and customers SalesForce has.
If you have experience enough to prevent downtime or to handle it when it happens, then perhaps that is better than risking that to someone else.
the 20 hours of downtime was probably BECAUSE of the amount of data/customers to handle. whereas if there is 1 customer, recovering isn't nearly as difficult or time consuming.
So what?
If your business is built around a system and that system goes down, would you rather that system serve billions of customers and go down for 20 hours, or serve only you and go down for 20 minutes?