Major data center power failure (again): Cloudflare Code Orange tested
blog.cloudflare.com
blog.cloudflare.com
Also — quite impressive to make major infrastructure and architecture changes in a few months. Not every organization can pull that off.
And have it work first time round.
Here is a free hint: By talking so much about where their data centers are located, on my view, they already failed item one on my check list. Principle number one of Physical Security is, you don't say where your Data Centers are, except of course to very restricted number of "need to know group".
As predictions have no value, unless they are made prior to events, I predict the next outage will be some common core component, with some on/off type of config, with some common core configuration, that "could not be foreseen". Then to the next blog...On to the next outage....
Would definitely be interested to see the detailed RCA on the power side of things. Not many people really think about Layer 0 on the stack.
https://www.edgeconnex.com/wp-content/uploads/2018/10/ECX-22...
It reminds me of the amazon guy discovering that there was no way to fail back power without an outage, then them going off and building their own equipment.
Anywhere I can read about that?
Pretty sensible IMHO - I live in a country with a reliable electricity grid, and outages due to UPS malfunction are about as common as power outages.
[1] https://www.datacenterdynamics.com/en/news/aws-develops-its-...
It doesn't take long without cooling to cook equipment to the point of failure or reduced reliability.
7 to 16kW per rack is common even in these older colo facilities.
And there never would have been enough UPS to make up for not enough replacement breakers on site.
I thought most DCs just pause the cooling until the generator comes up, rather than running the cooling on battery power?
They are expected to last that long, but if the batteries are on year 4 of their 5-year life, that may not happen. What also may not happen is the generator starting up.
Or the automatic transfer switch (ATS) not working properly: it should be on either input Feed A or Feed B, but when it tries to throw itself over (making a loud kah-chunk sound), it gets stuck in between—this happened to us once.
"The perversity of the universe always tends toward a maximum." — Finagle's Law of Dynamic Negatives
Fun story, small IT company had a natural gas generator installed at the new office they built out. Power went out for the first time, lights go out, generator kicks on, lights turn on, lights turn off. Long story short the electrician didn't have the right breaker size in the truck and used a small one. Building was empty when it was tested which is why bigger DCs have a load generator to test at load.
https://cloud.google.com/blog/products/compute/google-joins-....
In third party data centers you're co-located with other customers that have mostly comodity server hardware which won't have those options, they generally (at least in my experience) all provide UPS power as part of the facilitiy so you have it already and something like a fire if it did happen, would impact lots of other customers who likely don't have redunancy between data centers or fire compartments.
Additionally depending on your exact setup (it's possible you may have dark fibre directly out of the facilitiy, but often not) you're likely to be uplinked through powered equipment at the same facility, that wouldn't have the extra UPS power either.
Plus as others mentioned, the loss of cooling is the biggest problem. In fairness maybe that would buy you a few minutes for the kindof automatic switchover they talked about. But that's quite a bit of co-ordination to have your own power, to not rely on anyone else being powered and then to be able to co-ordinate your shutdown before the room gets too hot - and not be at risk of overheating the room for others.
Plus there are safety issues like emergency power out for fire fighters etc which again, I'm sure Google could deal with at a facility-wide level, but if you're a small fish in a much bigger data center its harder to co-ordinate that kindof thing.
So while I'm sure it's all possible, there are obvious headwinds in many directions.
IIRC, it's about 30 minutes—irrelevant of room size.
This is because while smaller rooms have less air volume as a 'buffer', they also you can't fit much gear. While huge data centres have huge volumes of air, they also have lots of gear.
So there's a coincidental proportional relationship what you (as a rule of thumb) have 30 minutes before things cook, regardless of room size. (At least for 'general purpose' computing: GPUs are probably another matter.)
The link someone else put about overriding faults might be it, if OP misremembered the problem.
And yet here we are, reading an article about "a total loss of power [...] following a reportedly simultaneous failure of four [...] switchboards serving all of Cloudflare’s cages. This meant both primary and redundant power paths were deactivated across the entire environment."
Hence why modern servers and their associated networking equipment have dual power supplies which could be connected to two seperate UPS systems. It would be very unlikely to have them both fail at once. In a less important home/small business scenario typically one supply is connected to the UPS and the other is connected to the wall via a surge protector.
Also, in my opinion, all UPS manufactures make horrible UPS equipment. There can be so many different types of failures or problems or glitches. You don’t want a UPS leaking or exploding. Also, there are laws in some states that only allow so many pounds of batteries in a building.
https://www.eaton.com/content/dam/eaton/products/electrical-...
I don’t see anything in the manual about needing to power it down for reconfiguration. All the relevant buttons, set screws, etc should be accessible without removing any dead fronts, so they can be safely accessed while the system is live.
Background note to HN readers:
Almost zero SaaS providers (or even CDNs) using the term "our datacenter" or showing their datacenters on maps etc. have their own datacenters. It's universal and normal. In general they have a server, a rack, a cage, in shared space, subject to others' policies and practices, and their neighbors.
This can adjust your mental model for accountability and your designs for resilience. You can even exploit this by colo-ing at the same addresses to get LAN latencies to your SaaS provider, CDN, or (sometimes) even cloud provider.
That HTTP API is backed by etcd which uses Raft and that is where running over a large area is more likely to cause problems. One approach would be to keep the etcd instances in one region (and probably also co-locate the scheduler, controller manager there), while having far-flung worker nodes. This creates a risk of losing the control plane but in most cases services would keep running (but would be unable to react to any further issues until the control plane recovers).
You would also want to carefully design your workloads with topology constraints and region-specific services to avoid high application layer latencies though.
Overall it's a fun thought experiment. In practical terms I think cross-datacenter in a small geographic area would work fine but I probably wouldn't want to run a single worldwide cluster, both for the reasons above and for other scaling reasons.
Is PDX still a single-point-of-failure for Cloudflare services?
It was 5-months ago [0], and if I understand the post - it sounds like it still is.
If anyone knows, I'd be curious to hear.
What this blog post talks about it is how this time nothing went down (or at least cut over within minutes due to presumably automated systems noticing & doing it) with the exception of analytics data which does have a single dependency and that’s determined to be “ok” (i.e. it’s an acceptable failure mode for that product & that product only).
There are additional failure complications required because PDX (& Europe) is composed of several independent data centers which aren’t supposed to all fail simultaneously. What’s pretty clear from the implied “we’re not happy either” is that Cloudflare isn’t pleased with their vendor’s separately located data centers still having correlated failures.
[1] https://blog.cloudflare.com/introducing-quicksilver-configur...
I hear you, but …
… what’s now core fabric of the internet (cloudflare) shouldn’t be suseptable to a single vendors data center going down to - knock out their service.
It create a situation where an attacker now knows all they need to do is target that single data center (PDX), cloudflare will go down, and now they can attack any website previously protected by cloudflare.
I’m not trying to be a hater.
I’m just trying to call out, it’s a bit unfair to solely point the figure at this data center provider when architecturally this issue shouldn’t exist in the first place.
Neither this outage nor the previous broke Cloudflare's protection.