Within 3 days another power outage at Linode (Fremont)
status.linode.com
status.linode.com
In their posting about the outage on the 20th they said: "At this point all we know is a severe lightning storm in the area caused a power outage and redundant UPS systems failed." http://status.linode.com/2010/11/possible-power-outage-in-fr...
Redundant UPS systems failed? And now it fails again? What kind of data center are they running in Fremont?
For example, power outage occurs at the same time the UPS batteries are being changed. Bypass fails and Diesel generators fail to kick in. Circuit breakers blow everywhere making it extremely difficult to get the generators back on line. This all happens in the middle of the night in a winter storm (or 'lightning storm) which causes a further delay in response time.
Been there, done that.
Edit: Also, bureaucracy, lack of documentation, and a manager CF added several hours to the outage. Sometimes you just have to STFU and let the geeks fix the problem.
I worked in a datacenter as well. The whole point is that you test those scenarios when it's not an emergency so you can be confident everything will operate smoothly.
However, as you point out it is far more likely that there are architecture or implementation issues that caused the repeated outages.
There are so many reasons that I don't want to have servers in California... high costs, dense population vulnerability to earthquakes and a hot terrorism target. Massive bandwidth but massive utilization, somewhat shaky power grid, heck even the state has financial issues. I picked Atlanta for my Linode server which is better (it's still a little close to the coast for me) and I host my WordPress sites at a company in Dallas, TX. Did you know texas has their own power grid - separate from the rest of the USA?
I personally didn't find this piece particularly great, but there's an entry in the genre for you.
I'm pretty sure there were Zombies roaming the aisles of servers where the power was out, that might help the box office sales.
That reminds me of hearing mention of "multiple six-sigma events" I think it was, when things were crashing in 2008: the only sane conclusion is that the models are broken.
HE doesn't have the best reputation for resilient datacenter services. However, redundant power is a very complex problem and prone to failure if you can't afford to do it right... which you can't if you're selling colocation for as cheap as HE does.
I really, really wish Linode would launch a sort of premium offering in better datacenters.
1) Many companies want facilities that are physically accessible to them, it's very difficult to sell to these companies if they'll have to travel 900 miles to do any work and don't trust the provider to do the work for them, or simply don't want to pay for it. Given the target market for hosting, Fremont's a pretty good option.
2) Depending on the service, a datacenter is only as useful as the transit providers you can access. We do all our hosting at a Chicago facility where we can get a cross connect to any bandwidth provider you can think of. Our services are also targeted at end users, so we're better off being centrally located and using providers with extensive peering agreements (less latency, woo!). If we had serious computation requirements rather than delivery requirements, one of those spiff Icelandic datacenters might be a good choice. I'm unsure of what Google and others stick in their rural datacenters, but I doubt it's anything that slows down your experience.
It is probably going to be almost impossible to have a "carrier neutral" hotel in rural Oregon but you might get a single ATT or Level3 or if you're really sucking it, a Cogent line.
HE Fremont is specifically great for reaching markets in APAC and HE itself is pushing IPv6 quite hard and so is one of the few providers that make it widely available.
Fair enough, but there's plenty of middle ground between Prineville and the Bay Area. The latter is expensive - it makes sense to have the really high-end things there, like Google R&D, not commodities like data centers and factories.
IMO. I'm certainly not an expert in that sort of thing though.
Incidentally, as a native of Oregon, I still think it's pretty funny that Facebook is building a center in Prineville, heretofore best known as the home of Les Schwab "Free Beef!" Tires.
But you are definitely right that most people don't need that kind of silliness.
I came from North Carolina where Google and Apple saw the lower power/employee costs as a reason to open datacenters. In the western part of the state is a beautiful city called Asheville which was never known as a network hub. But because of the geographic location halfway between Atlanta and DC a company(uberbandwidth.com/netriplex) decided to create a datacenter there. Lower costs...but the only way IN and OUT of their network was a backhaul through Atlanta or through DC. If one of those backhaul lines go down they've basically lost their single selling point.
I wouldn't so much wish for a premium offering, just better datacenters in general for Linode's racks.
It could have been worse in Fremont, I suppose, or that might just be a bit of puffery being used to explain the outage.
i colo'd at HE until last year when they ran ~400V through my racks which had 110V circuits, several PDUs and servers were damaged. that was the cherry on top of the sundae though, there had been 2 power outages at that point and 2 other outages after that point.
Plan for a minimum of one power outage every two-three years and you won't be disappointed.
I feel for the HE guys - back to back power outages has got to be killing them right now.
Dallas - 99.951% Newark - 99.969% London - 99.986% Fremont - 99.989% Atlanta - 99.995%
http://www.365main.com/status_update.html
The moral of the story for me is;
* These things are complicated
* Failures will happen
* You have to be prepared to deal with them
Back in 2008 an HE based colo "McColo" were shut down because they were hosting a HUGE amount of botnet controllers, spamming operations, and similar shady operations. When they were shut off some security firms saw a 50% drop in spam going through their firewalls.
However it only happened after immense pressure on HE and other providers involved by Google, Washington Post and all sorts of other players.
HE would have been aware because large IP blocks were being blacklisted (I heard at one point all of HE Freemont's IPs were blocked by some of the more extreme SBL lists) but they turned a blind eye and/or claimed Ts & Cs were not being infringed.
More here http://news.cnet.com/8301-1009_3-10095730-83.html
I found that highly irresponsible, both in terms of the detriment to their other colo customers who were sharing the BGP-level bandwidth but also from a wider 'being a good actor' perspective.
Given that they are also on a fault line and that good connectivity to Europe is more important to me than Asia, I would prefer to host on East Coast.
I'm actually in Linode's New Jersey and London Colos, and they are both excellent.
I would actually call pulling the plug without due process irresponsible.
Also, I question your claim about it happening only after "immense pressure." Your own link praises them for their response, and some googling suggests similar wording in all coverage I can find.
http://news.ycombinator.com/item?id=1926368
Actual IRC info...
Server: irc.oftc.net Channel: #linode
AWS is a cloud product, with pros and cons -- instances can die at any time and data (RAM and disk) won't persist, network is slightly strange with their NAT IPs but in return you get a setup that lets you connect with big storage (S3), potentially large clusters (the new GPU core product) etc.
Linode is a VPS which is just presenting you with an abstracted server. If the instance or the hypervisor gets bounced anything on disk is preserved. Networking is normal (standard IP address) and everything runs as if it is a bare metal server (more true for XEN based VPS's like Linode rather than SolusVM)
Search: 0.1% of 1 month in hours
Result: 0.1% of (1 month) = 0.730484398 hours