OVH Incident in Strasbourg
status.ovh.com
status.ovh.com
and on https://twitter.com/ovh_support_en
"SBG: ERDF is trying to find out the default. 2 separated 20kV lines are down. We are trying to restart 2 generators A+B for SBG1/SG4. 2 others generators A+B work in SBG2. 1 routing room is in SBG1, the second in SBG2. Both are down. "
"An incident is ongoing impacting our network. We are all on the problem. Sorry for the inconvenience."
"SBG: 1 gen restarted."
"RBX: all optical links 100G from RBX to TH2, GSW, LDN, BRU, FRA, AMS are down."
(edit: parent's post was edited after I posted)
Pro-tip: self-hosting status page is maybe not the best idea.
[1] https://twitter.com/awscloud/status/836656664635846656?lang=...
Lister: What's the damage Hol?
Holly: I don't know Dave. The damage report machine has been damaged.
SBG going down due to a quadruple power failure (both grid connections and both generators) is quite spectacular.
Today, I'm glad to have moved away all my production environments as well.
That's just the problem with services as AWS, traffic is too costly to have production live with another provider as well.
Having one AWS account scares the crap out of me as well. It’s never a good thing if all your eggs are in one basket.
My money is on stuff spread across Bytemark, Linode and DigitalOcean with a DR plan involving mostly automatic recovery.
AWS doesn’t get a look in as it is extremely costly to port away from anything that isn’t bare metal and pipes.
Ansible works for this stuff as it allows you to can the task of “quick get me a production environment up on Linode!”
If you can afford some downtime you don’t need a hot standby just roll out everything into new provider and you’re done.
I’ve done this on very large scale environments and small ones and it’s achievable for even small organisations. The killer is avoid anything you can’t run on bare metal servers.
1) your DNS provider having issues (even Route53 sometimes has them, https://mwork.io/2017/03/14/aws-route53-dns-outage-impacts-l...)
2) legal issues, when one of your domains gets seized or the provider gets pressured into cancelling, just as has happened with Pirate bay, SciHub and friends, gambling or sites with user generated content that may be illegal or frowned upon in some countries or for the latter to the Nazi site Daily Stormer (although I'm glad for it being down, it's a perfect example what can happen in a very short time frame)
3) (edit, after suggestion below) the entire TLD going down because the TLD DNS provider has issues, which also happens from time to time.
Just fire up another one, install everything on there, switch off the old one and forget about it. This means never paying for more than one month in advance.
Servers break. Providers go down.
It happens to all of them. Whenever I read comments like yours I wonder where you moved your servers to, and if you'll move them again when that goes down too.
Same goes for backups: you don't have backups until you tried to restore them.
Is there not a way to test generators without actually having them power the live datacenter infrastructure? I mean, simulate the exact generation and load requirements that the generators will face?
I don't know if it's feasible to dump all that power to ground or whatever, but that way you could test the generator under full load at will and identify issues without impacting the (operationally) live datacenter itself.
This equates a bit in my head with the 'verifying the backup' bit you might get in software, whereas actually using the generator to power the live operational datacenter equipment would be more like 'restore from backup'.
I don't know if it's possible though.
You know all that cooling equipment data centers tend to have?
Dumping power to ground is also known as an electric furnace.
Now you have twice as much heat to move, but you only have the usual cooling system. Toasty servers are sad servers. Toasty engineers are dead engineers.
I don't understand why you'd be putting a generator test load through the servers in addition to the normal power supply??? Why would you do that???
I would have thought that you have normal operation going on in the datacentre, normal cooling infrastructure, normal power coming in. Then the generators turned on with their in built cooling, then them delivering power to the test area and not the datacenter - which is then the only additional area you'd need to cool.
Seems like dealing with the heat would be not that hard, but perhaps I have too much faith in engineers? :-)
Even just a massive heating element in a tank of water, giant kettle style would work wouldn't it? Big kettle mind you, but big tank too. Seems like a cheap way to test the generator at full load for an extended test?
That would exercise transfer switches in addition to generators. Transfer switches are always energized, except for the time an equivalent of a big red mechanical switch is flipped into "OFF" position. When it is in the off position, neither main or generator are going to provide power to the customer. The biggest power consumer is actually a cooling system. If a HVAC system stops functioning in a typical data center, the temperature would quickly rise to the level of Arizona desert, destroying metric tons of equipment. Transfer switches have certain properties where after a flip they may not go back to the correct state on power restore ( say 1% ). Normally it is not a big deal because you really need to have lose power under full load often to get bit by it and if you are losing power that much you should probably address it with the utility company. However, if you are doing your full test once a week, in one year you would introduce 52 power failures.
Say your transfer switch is stuck in a wrong position. Now you need to shutdown all heat generating equipment to when you drop power to fix/replace the failed component you don't melt your data center... Congratulations, your data center now has to go offline.
I accompanied my dad (power engineer) to a water purification plant where they were testing new equipment for the back up generator. There their weekly tests involved moving the entire plant to the diesel generator and running it of back up power for a couple of hours (once you start a big generator you have to let it run or it wont last long).
Potential problems for your generator that a resistor bank wont capture include, power factor (phase shift from a motor or switching supply), harmonics (from switching power), startup transients (from every power supplies' capacitors).
All these things can trip the generator, or worse, burn it out.
So if you can't test with the real load, supersize it!
P.S. Every test is a simulation of reality. At Fukushima the diesel generators flooded. Lesson - the unknown reason that'll knock out your grid can knock out your backup
P.P.S If you can, gently turning the load back on is very beneficial. Don't flip the master switch that controls all your load - flip a part of your load, wait a while for the system to stabilize, and flip part of it back.
well the lesson there was more like, that it's stupid to put your diesel generators deep in the ground when they should sustain burst sea level raises. (well I think it's never a good idea to do that, I've seen special places to put them even deep inside germany, just because some panicful people that might think that it still could overflow with ground water, etc)
The basement was, according to plan, protected by a seawall.
The height of the seawall failed to take into account the fact that on a subduction tectonic plate, as pressure builds, the land-side plate rises, and as the earthquake relieving that pressure strikes, the land falls -- by as much as several meters.
The seawall's height failed to account for this.
That among other elements, but it proved sufficient to kick off the Fukushima disaster, given other aspects.
In hindsight, placing the generators at ground level in an elevated location might have been a better bet. Or locating the entire generating plant further upslope.
yeah, well thats what I meant. I mean they were buried deep... There were even studies that this was dumb:
- https://news.usc.edu/86362/fukushima-disaster-was-preventabl...
- https://en.wikipedia.org/wiki/Fukushima_Daiichi_Nuclear_Powe... (section end)
- http://carnegieendowment.org/2012/03/06/why-fukushima-was-pr...
and basically tepco knew that. they were just too lazy (it was prolly too cost intensive) to do something against it (i.e. placing them higher or creating water bunkers (u-boot style)..
besides the generators there was also the human failure part. (most of the time human failure happens. I mean If I do not automate something, I might do it right 3 times and the fourth time I most often fail hard...)
For example you switch off the data centre circuit breakers and everything fails over to generators just fine. Test successful, right?
Then when there's a real outage you have problems because the operations team's computers have all gone off, so they can't migrate load to a different data centre. It didn't happen in testing, because they aren't in the data centre so their kit isn't connected to the breakers you turned off.
Or it turns out the wireless APs aren't on UPSes. Or it turns out there's a switch in a closet somewhere that isn't on a UPS. Or they tested for a single loss of power, but when the mains power toggles on and off every 30 seconds the UPS batteries get run down. Or they need to top up the generator and they discover you can't get fuel delivered at 9pm on a Friday. Or the generator doesn't recharge the UPS, but you have to turn off the generators to refuel them. Or a guy had a standalone UPS for his desktop, but his monitor wasn't connected as the UPS only came with IEC C13 power cables and his monitor needed IEC C5...
But you want to test the things you haven't thought about.
There's tooling available and running at some of the larger companies (like Netflix) called Chaos Monkey, which triggers random outages constantly to determine if the system is resilient and self-healing.
EDIT: He did indeed delete the tweet, [1] is the url in case anyone knows a website archiving them quickly enough.
[1] https://twitter.com/olesovhcom/status/928536233311076353
Usually.
Unless by stronger you mean some kind of a change in mentality over care for the service, but that would probably have a limited lifespan until it's back to normal.
If a service never fails then I don't know how well they can recover. I don't know anything about their failure mode.
In this case, I'm learning how OVH handles failure modes, how well they handle it, etc.
I can observe how they will treat such things in the future.
Thankfully many times we're reminded that there are good people out there working hard against difficult constraints and they finally get their chance to do things 'correctly' in the wake of the SHTF.
As a more technical user it's nice to have providers that give this information rather than the boilerplate "Issue with an upstream provider" over and over.
Just in case you wonder why your sites don't work, even if you host them somewhere else.
I paste the report so far:
-------------
FS#15162 — SBG
Attached to Project— Network
Task Type: Incident
Category: Strasbourg
Status: In progress
Percent Complete: 0%
Details
We are experiencing an electrical outage on Strasbourg site.
We are investigating.
Comments (2)
Comment by OVH - Thursday, 09 November 2017, 10:55AM
SBG: ERDF repared 1 line 20KV. the second is still down. All Gens are UP. 2 routing rooms coming UP. SBG2 will be UP in 15-20min (boot time). SBG1/SBG4: 1h-2h
Comment by OVH - Thursday, 09 November 2017, 12:04PM
Traffic is getting back up. About 30% of the IP are now UP and running.
-------------
VPSes are still marked as read in the dashboard. I can't access mine.
Comment by OVH - Thursday, 09 November 2017, 12:44PM
Everything is back up electrically. We are checking that everything is OK and we are identifying still impacted services/customers.
Comment by OVH - Thursday, 09 November 2017, 13:25PM
Hello, Two pieces of information,
This morning we had 2 separate incidents that have nothing to do with each other. The first incident impacted our Strasbourg site (SBG) and the 2nd Roubaix (RBX). In SBG we have 3 datacentres in operation and 1 under construction. In RBX, we have 7 datacentres in operation.
SBG: In SBG we had an electrical problem. Power has been restored and services are being restarted. Some customers are UP and others not yet. If your service is not UP yet, the recovery time is between 5 minutes and 3-4 hours. Our monitoring system allows us to know which customers are still impacted and we are working to fix it.
RBX: We had a problem on the optical network that allows RBX to be connected with the interconnection points we have in Paris, Frankfurt, Amsterdam, London, Brussels. The origin of the problem is a software bug on the optical equipment, which caused the configuration to be lost and the connection to be cut from our site in RBX. We handed over the backup of the software configuration as soon as we diagnosed the source of the problem and the DC can be reached again. The incident on RBX is fixed. With the manufacturer, we are looking for the origin of the software bug and also looking to avoid this kind of critical incident.
We are in the process of retrieving the details to provide you with information on the SBG recovery time for all services/customers. Also, we will give all the technical details on the origin of these 2 incidents.
We are sincerely sorry. We have just experienced 2 simultaneous and independent events that impacted all RBX customers between 8:15 am abd 10:37 am and all SBG customers between 7:15 am and 11:15 am. We are still working on customers who are not UP yet in SBG. Best, Octave
Fix (debian-like):
sudo apt-get install bind9
Then put in /etc/resolv.conf, if it's not already there: nameserver 127.0.1.1
This runs a local nameserver that you use directly for resolving.Oh, obviously, you need resolving to install the resolver :) Hope you have a 4g connection available.
Alternatively, you can just use google dns:
nameserver 8.8.8.8
nameserver 8.8.4.4[1] http://status.ovh.net/?do=details&id=15162&PHPSESSID=7220be2...
https://twitter.com/olesovhcom/status/928541667283623936
EDIT: Maybe related to the Cisco issue?
https://blogs.cisco.com/security/cisco-psirt-mitigating-and-...
Their explanation will be really interesting. All those data centres are pretty useless if they have a single point of failure.
One datacenter suffered massive power failure while the other hand network equipment fail at roughly the same time.
According to them, they do their routing in SBG, so it's plausible that it could lead to all of their network being down.
Apparently, the root cause of that issue is a critical software bug in Cisco NCS 2000 transponders.
Their interfaces lost their configuration, and they re-applied configuration, and state came back. This does not equal critical software bug.
> "One of the solutions is to create 2 optical node systems instead of one. 2 systems, that means 2 databases and so in case of loss of configuration, only one system is down. If 50% of the links go through one of the systems, today we would have lost 50% of the capacity but not 100% of links."
This is a crap mitigation. They're still depending on the same hardware and process that led to the first outage, only now there's more of it, so there's more chances to fail.
If they had continuous configuration automation they would have detected when the router's state changed, identified the missing bits, and applied configuration.
"New" routers (as in, since 2011) have APIs and can even run code directly on the router in order to fulfill these requirements. Cisco has multiple white papers, and even provides complete products to manage and certify configuration is applied as desired, even in cloud-agnostic multi-tier networks. Even on old routers, practically all config management solutions out there have plugins to manage Cisco routers.
It's also ridiculous that they had no access to remote hands. This is IT 101.
Maybe their main DCs, or their largest, but not all of them. I have virtual servers in thier Quebec DC (BHS) and it hasn't gone down since the last time I rebooted it.
http://web.mit.edu/2.75/resources/random/How%20Complex%20Sys...
ovh.com looks down for me too.
You can check it's hosted by OVH:
$ whois $(dig sudokugarden.de +short)
Does this happen often with OVH?
EDIT: Hosted in EU.
Anybody knows ETA?
My OVH servers in France are all inaccessible.
Opened up Age of Empires II....no connection. Go to website for game servers..."Our provider, OVH, is down...."
Go figure.
My data is with me in Europe and the company that has my data is in Europe with me too. If I was using DO or AWS, then my data may be in Europe but the control over the data is in the US, free for all to the three letter agencies and lacking privacy laws there.
Sure, it's less dramatic than AWS going down, but it still hits, and hard.
"guys we have a single point of failure in our architecture with SBG, maybe we should...
- naaah it's fine, we do not have time nor resources"
Then shit happens.
edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me (or backup systems, like generators, not being correctly tested... whatever). I am really curious about the future explanation with what happened exactly
edit2: Why all the downvotes? Even the status page of OVH is down, do not tell me it is good design. We are not here to be charitable, but realist.
No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can happen with one product (eg S3 failure) but datacentre switches should work even if the rest is on fire.
So, that one datacenter caused all other datacenters to die..?
It's explained here https://twitter.com/olesovhcom/status/928587258583748609 in French, they might post an English-language translation soon.
They just announced that they want to be a "hypercloud" provider on the scale of AWS and Google Cloud. I really hope that a power failure in Virginia couldn't bring down all of AWS.
I can not imagine they weren't tested.
But even the most rigorous testing can never reduce the total failure risk to 0. It seems OVH just got very very unlucky.
So our lead Ops guy is unplugging half the redundant power supplies and plugging them into the new circuits, but a few critical servers are single PSU still, on UPS units. This is how he discovers that one of our PSUs has rotted and only has about five seconds of reserve power in it. Big outage, no bueno.
I'm much more interested in why the failure cascaded to other data centres – that's exactly what shouldn't happen.
While generators are usually able to handle this with sufficiently low failure risk, the risk is increase due to the changing temperature
[1] https://www.accuweather.com/en/fr/strasbourg/131836/daily-we...
Their generators are almost certainly tested frequently. But there could be any number of causes underlying the failure, and unfortunately sometimes failure does happen.
There was a DC in California a few years back that had three generators fail during a scheduled test and one of their server rooms had a blackout because of it.
It happened to GCE last year[1], though it only lasted 18 minutes.
[1]: https://status.cloud.google.com/incident/compute/16007?post-...
but
support/administration does not work well, i have a lot of really weird story's with them, from them plugging in a keyboard in our server to reboot it (without any reason) to taking down a server for a requested maintenance only to notice after 4 hours of downtime that they did not ask their bosses if they were allowed to even perform the maintenance requested (and then not getting permission to do so after another 2 hours ..)
For me it feels like there are some really deep issues somewhere in the whole administration that make incidents like this no real surprise
Problem is most other providers dont work any better, so ...
Everyone makes mistakes, let's just hope they learn from it.
The prices are high? Compared to what? Cheap is their raison d'être.
Sry if that caused confusion
Just the bandwidth alone would cost us 3x what we pay at OVH if we were with one of the big cloud providers.
Yup, been there multiple times in smaller hosting companies.
It's basically how it goes. They don't get serious about outages until revenue is severely affected and the brand damaged, they don't get serious about security until there's been a big breach or sales are lost because of lack of certification.
> in smaller hosting companies [like ovh]
> in smaller hosting companies [than ovh]
> Yup, been there multiple times, in smaller hosting companies.
I find it funny that anyone would think I would or could refer to them as small (and get away with it)
EDIT: The title on HN in misleading, summary from their CEO here – https://twitter.com/olesovhcom/status/928592231807713280
It is rather counter-intuitive. In neteng and dcops, "I don't know and I cannot find out. I can only attempt to mitigate what i think might have caused it for next time" is a very reasonable answer to 99% of the "why this happened?" questions because in order to replicate the situation to test the theory one needs to recreate the same problem again on the same scale.
This also means that certain things cannot be tested. Most of generator tests are garbage - turning on generator and running it without production load delivered over the transfer switch does not test anything other than that one can turn on a generator and run it. The problem typically happens not because the generator ( also is there the generator or the first and the second generator? Why is there no generator bank for a non monkey-sized company? ) does not start - the problem is because over time transfer switch develops a problem and unlike generators it is not possible to test a transfer switch where in the event of a test failure the customers won't lose power unless the data center is designed from the beginning to deliver A and B powers over separate circuits to every single customer and every single customer has per system ( not per rack ) transfer switches.
Of course it costs a lot more money, something that companies are reluctant to spend.
If someone doesn't know why a failure occured and they can't find out then they aren't looking hard enough
(In one case I experienced, the vendor reassured me that what I reported was not even technically possible -- until one of their engineers flew out and witnessed it firsthand.)
If your CD pipeline deploys corrupt apps, even your production database can get compromised,forcing a full restore. No matter how long that might take
Thankfully,I didn't have to witness that yet.
@DNS_BORAT https://twitter.com/DNS_BORAT
@InfoSecBorat https://twitter.com/InfoSecBorat
@KanbanBorat https://twitter.com/KanbanBorat
@mysqlborat https://twitter.com/mysqlborat
@NetEng_Borat https://twitter.com/NetEng_Borat
@secure_borat https://twitter.com/secure_borat
@SecurityBorat https://twitter.com/SecurityBorat
@Sysadm_Borat https://twitter.com/Sysadm_Borat
06:15 UTC SBG serves failed.
OVH network weathermap: http://weathermap.ovh.net
Btw. First post: https://news.ycombinator.com/item?id=15660524
My submit timestamp: 07:21:25
Your submit timestamp: 07:28:23
About the title. I did consider the title for a while, because I wasn't sure how bad the situation was. But from my own independent monitoring system I did see that RBX and SBG servers were unavailable. Of course I also did some basic trouble shooting and confirmation work before posting.
Btw. Right now, there's some network traffic present on SBG network. Let's hope that the systems are soon up'n'running.
Yeah, I expected this to clear up in a matter of minutes.
Now it seems to be a shitstorm of historic proportions...
Even if they didn't, clearly many services are up and running normally, so saying "all datacenters are down" is just a lie.
I think we're now going to have to look into multi-provider options. The only way to be solidly up is to be hosted by more than one company at more than one data center.
I've also heard stories of billing nightmares where you get locked out of a cloud provider account, so that's another thing.
I guess this is already a reason by its own. It, among other problems, is what happens when we go from small "local" providers you can actually call to automated global providers that cannot provide immediate support even if they tried.