Verizon has no excuse for its planned cloud outage
infoworld.com
infoworld.com
Is it somehow related to their main business? Do the servers run deep within their network allowing quicker access for mobile users? What is their unique selling proposition?
They have a whole "why Verizon" section[1]. I don't feel like it answers this question.
Wouldn't be such a problem if they did a good job at their core business but as we learn more about google and Amazon's network, it seems like the telcos may fail there too.
The point of the Cloud is not that a server never restarts or has downtime. It's that your app runs on many servers in different AZs and regions such that any one server failing is not going to have a real impact and can be easily and quickly replaced.
AWS, when bouncing their servers, did them one az and one region at a time. Because of that, if you ran across multiple AZs you would get to test your failure scenarios, but not be down.
The author's focus on other cloud providers having restarts makes it sound like he's saying "Verizon is bad, but in bad company". There is no comparison to be made between taking down your whole cloud, and losing a server here and there.
All of the bigger Clouds (Amazon, Rackspace, SoftLayer) handle their security related restarts this way, one area/region/DC at a time carefully making sure that well designed systems don't notice downtime.
What Verizon does is a totally different thing and something most "important but not life threatening" systems do not design for: Taking down the whole cloud with all servers at the same time.
With that said, I have zero hope that most enterprises will be able to get a decently solid failover strategy for their cloud applications because an awful lot of them are barely able to get anything to even one cloud provider setup at the most basic layout. Half the folks I'm aware of don't even have cross AZ redundancy let alone cross region failover. When people are turning off 3-4 VMs to save money, you're in no position to think about cross provider load balancing.
Most of the enterprises I'm familiar with would throw Verizon more and more money in the blind hope that it'll make their service more reliable. They'll do this to an incredible dollar amount partly because it'll still be cheaper than their old, traditional in house IT that cost multiple orders magnitude more but got way worse availability. Stuff like this scheduled maintenance shows how foolish it is to believe that someone will magically become ultra reliable by you virtue of you paying them to be.
It doesn't look like customers were given much warning either. This story was originally published on the 6th of Jan. Could you imagine trying to find alternate hosting setup by the weekend if you have any kind of availability expectations? It seems like madness to me. Even if you did move yourself to another host to cover this 48 hours of downtime how likely are you to move the majority of your business over to AWS, Google Cloud, Azure etc.
The lack of notice on this seems to be a bigger issue to me than the fact that Verizon is taking their whole cloud out of service for 48 straight hours.
Then again, I don't know why you'd think Verizon was a good hosting provider in the first place.
This is a lesson about the difference between hosting your own apps and hosting customers. All internal Google apps are designed to survive failure of an entire datacenter, so it sounds like they reuse that mechanism for maintenance rather than the riskier practice of maintaining a datacenter while it's live.
Getting back to Verizon, I work in an "enterprise" and scheduled multi-day datacenter maintenance tends to happen about once a year. Many of Verizon's customers would probably have no problem taking down their infrastructure for a weekend, but they feel helpless when someone else does it to them.
Not to mention that if you buy in to certain services, say, any AWS architecture beyond EC2, failover to another provider becomes a lot closer to impractical if not impossible.
Two hours is sufferable...
Natural (and unnatural) disasters happen, and you have to be prepared for this contingency. There are many tools out there which can help with this, but a competent sysadmin will help a hundredfold ensure that your business continues when the unthinkable happens.
Heck, only today one of Amazon's new datacenters on the east coast had a 3 alarm fire[1]. Didn't end up having any instances fail, but we did notice some of our services have problems that coincided with the fire, and we were ready to fail over the moment instances started dropping.
[1] http://money.cnn.com/2015/01/09/technology/amazon-data-cente...
Yes. I say this all the time and continue to be amazed that this is such a foreign idea to many cloud users. When you stick to the infrastructure-as-a-service offerings, like EC2, and steer clear of cloud vendors' proprietary platform services, then you evade costly vendor lock-in. And this comes with very little added engineering cost.
Granted, each organization will value vendor independence differently, but I suspect many organizations don't give enough consideration to worst-case scenarios.
Hey, let's try a little experiment on the subject.
I'll sit here and do nothing starting right now, simulating unscheduled downtime, for the next two minutes.
You sit there and do nothing starting right now, simulating unscheduled downtime, for the next two days.
ETA: two days downtime divided by 365 days a year is over 15 minutes per day, or almost 2 hours per week ... any provider down that much & often would lose customers fast.
The culture of technology rightfully abhors downtime, and that's a good thing because it's our job to keep things up. But in reality there are very few companies who cannot survive 2 days of downtime. Sony Pictures is still in business, for example, despite an effective downtime of weeks.
Verizon telling you that some random servers are going to be down for two days admittedly might be worse than taking one of your own systems down for a maintenance window is, especially if you didn't expect in advance that such a thing was going to be a possibility, and made the mistake of putting really-can't-be-down stuff on Verizon's cloud.
Does Verizon not work in IT for a living? Why does their architecture not allow them to switch over to a backup despite having many days of warning?
It is slightly more complicated than that.
> There should be no cloud outages, ever. Got that?
Really? That's not how the real world works. This "article" is pretty crap.
Nobody expects 100% uptime. They also don't expect 100% down for days.
Also, nearly nothing will give you 100% uptime. That's fault tolerance and almost nobody has it.