Cooling related failure (in Google London DC)
status.cloud.google.com
status.cloud.google.com
Obviously the weather is easy to blame, but I wonder if the underlying cause is the same datacenter. It's kind of annoying when clouds use availability zones as such an opaque thing, it's not possible to map a zone from one cloud to another and potentially your failure domains overlap.
Shameless self promotion: I have a map of all the cloud regions (but not going into as much detail as availability zones): https://cloud-regions.bodge.cloud -- the clouds just don't publish the data.
It's also not appropriate in all cases to spread out, for example the nearest zones to London are around ~7ms away in mainland Europe.
I just plug into the wall socket and it works.
How could I choose to connect things to L2?
Because phase 1 is picked so often, they actually rotate the phases at the street ie switch 1 2 3 with 3 1 2 to avoid having ie 300% more load on phase 1 and having the grid just die.
AFAIK phase order matters for e.g. three-phase motors, they will spin in the wrong direction if wired with the phases in the wrong order.
It seems blue is strongly favored by psychology students for example : https://www.livescience.com/34105-favorite-colors.html
The GP is also a bit misleading. Most residential properties in the UK are only fed one phase - either the street only has one phase or the grid connects houses up L1, L2, L3 in sequence. Of course I'm speak about detached and semi-detached houses, in larger groups of houses (e.g. blocks of flats) you usually get 3 phases, and one of the jobs of the electrician is phase balancing.
https://www.newfound-energy.co.uk/electrical-three-phase-wir...
That's bollocks or speaks for a completely incompetent electrician. Usually, you have three phase rails coming out of your GFCI that distribute the phases sequentially to each breaker [1] to avoid that scenario, as well as ending up with a tripped main breaker because the customer overloads the phase.
The only case where great care is being taken to avoid random assignment of phases is in event stage technology - you do not want your lighting dimmer packs on the same phase as your amplifiers because dimmer packs inject extreme amounts of EM noise, and you also have to take care to not put overloads on any single phase anywhere. However, it's questionable on how long this will stay relevant, given that most lighting load is moving off to LEDs.
[1] https://www.amazon.de/-/en/Busch-Jaeger-Hage-Phase-KDN380A-3...
What you describe is... beyond awful and an absurd fire risk.
They are very slowly modernising here, but they seem to have a cultural resistance to change/modernisation; they're in love with the old days.
Very different to Aus/NZ where we're eager to get the latest stuff/NZ is often used as a test market.
AWS at least in the past use to co-locate with other third party servers in the data centre. And so you if you had such a server you could ping AWS endpoints to triangulate where physically those servers might be.
If yesterday's temps lasted for a week instead of a day or two a lot of essential stuff would stop working.
I was trying to point out that high-end hardware should survive a couple of days despite a data centre losing its cooling. it is designed for exactly that situation.
You are right, and it's arguably harder for those smaller boxes sitting in cabinets / on telegraph poles, despite being much lower power systems to start with. It might be 50C+ in there before they even turn on! Those things might have a two-stage boot process where they just run their fans for a bit to cool things down before actually booting the main system. It must be a real nightmare for entirely passive stuff, I have no experience with that.
Sounds like the downside of the "pets versus cattle" methodology is that natural phenomenon will wipe out your entire herd, while the pets survive.
Being able to look up at the buildings there and know there are indeed 5 different clouds somewhere up above my head in that specific location would be really cool. Being able to point at a specific building would be even cooler.
I do of course (sadly) appreciate the flip side of this coin which is one of (many of) the reasons precise data is not published. So I guess I'm just wondering out loud, probably rhetorically :), how I might find out one day. "<-- That building" is enough resolution for me :)
EDIT: A quick Google found baxtel.com (among other websites) has address-level locations for most providers. The buildings are all so boring, haha! (Understandably so though.)
Secondly, within a limit of my understanding, the efficiency of a turbine system hot-to-cold is affected by the climate it operates in. The ambient temperature, humidity and pressure affects the final stage.
Both things might mean that in times of high heat and humidity, the electricity supply system is least able to cope with increased demands for cooling systems, which will themselves draw more power fighting the weather.
Separately the HVAC systems for the DCs will have been designed for a specific climate, with margins. I guess the sustained change in night and daytime temps and humidity has hurt their efficiency too, in this window of time. They'll be fine when the weather system passes through, as will the supply network.
Met Office says both overnight and daytime peak temps for the inland south have been records. Thats where a lot of ICT infrastructure is.
IE: less of that relieving overnight low.
Usually this is only partially true. Systems have enough margin to account for such conditions (cooling pumps have to pump more water to compensate for higher temps). Also to note that the output of a turbine can be kept constant. Only the efficiency will come down slightly.
Exactly right. For a Tier IV data centre that margin is ASHRAE N=20. For London City until recently that was only 34.5C (dry bulb).
It's cathartic seeing the capitalists lack of foresight burn them time again, but also sad, knowing the response to this will be just buying slightly better HVAC and other systems, which will be globally sourced no doubt and burn carbon to produce and probably damage some ecosystems in the resource gathering process. Then when temperatures get even hotter its time to buy new systems etc.
Once again capitalism favors this outcome, because it allows multiple sales opportunities over time with the potential to increase monetization at every one, versus buying one system that can handle wider temperature swings and being done with it potentially forever if its made to be repairable/upgradeable. Even if you invented such a thing and sold it out of your garage, GE or whoever your competitors would be would ensure you have difficulty sourcing necessary parts or coming to market or getting word out to your potential customer base. Investors you need to afford to grow under these conditions would expect you to play the game and start cutting costs and engage in the rat race, or be replaced.
It's going to be hard to work our way out of climate change with the degenerate, consumptive nature of capitalism mining the planet while sucking resources from the wider economy where this salvation is to be invented, to the top that will always be able to afford to hole up and hide away from whatever disasters befall the working people.
Your whole premise ignores the billions of people who happily use electricity and fuel everyday and happily ignore the consequences. These same people then (literally) riot when prices go up by a relatively small amount.
People are greedy. They want more for less. We can disagree about the best form of economic association, but it’s disingenuous to ignore the reality that climate change is driven by the masses.
And the answer to a "tragedy of the commons" scenario is regulation: either intervene directly by banning undesired behaviors (e.g. ban flights on routes that are served by rail, such as France does) or tax it to make undesired behaviors unprofitable.
The problem with the latter is that you will always have people rich enough to simply pay whatever tax is asked and social resentment will grow as a result ("why should the lower classes bear the load of climate change and the rich enjoy three-minute flights to save a 40 minute road trip [1]?").
[1] https://www.buzzfeednews.com/article/stephaniesoteriou/kylie...
Lack of foresight isn't exclusive to capitalism. Would socialist data centres be built to handle temperatures several degrees above the highest ever recorded temperature?
Even in a socialist economy, building systems with an excessively high tolerances would be seen as a poor allocation of resources.
It could be equally argued that capitalism favours private companies like Google ensuring their DCs are as fault tolerant as possible, to ensure they have a competitive advantage. There's also plenty of cases where companies sell unnecessary and excessive goods and services to maximise their own profit.
> and being done with it potentially forever if its made to be repairable/upgradeable
DC cooling systems are repairable and upgradable. They're a far cry from a residential split system AC unit.
Socialism can prioritise quick fixes or long term solutions, but it depends on the wisdom of the people involved as to whether they'll build in enough capacity to allow for future climate changes.
Huh? GDR and Soviet made machinery, vehicles, even glassware for pubs [1] was made with sometimes ridiculous margins and tolerances to ensure longevity and easy repairability and was famous for it. Even pre-reunification Western made products such as Bosch, AEG, Hilti or Siemens were famous for building stuff that sometimes outlasted the owners (such as my 80s Hilti drill, which served three generations of my family and likely will still work when I have children of my own).
Bollocks. Soviet machinery is simple and crude, typically many decades behind their western counterparts.
"Keep it simple and stupid" is a tried and true engineering principle. The higher tolerances a design uses, the easier it is to manufacture and to repair, and the less likely it is to fail from wear and tear in the first place.
A socialist economy, where waiting for a new car could take anywhere from five to twenty years (!), definitely has to prioritize simpler, more (fault-)tolerant designs even if that takes a bit more resources to account for said tolerance. For example, a modern car heavily using fibreglass and plastic in the chassis may weigh a good load less than your average Lada or whatever that was made out of metal, but it could easily be repaired by your average farmer using tools they had in their shed.
Random side fact, this is a major cause why farmers are paying record prices for tractors nearly half a century old [1]. Or why the Russians are currently using so much ages-old stuff in the Ukraine war - modern tanks require a lot of logistics for repair and spare parts, but these old Russian clunkers? You can piss into the tank and it will probably drive on it. (Yes, I know, the Russians haven't been maintaining their tanks properly, which is a major factor in why they were not able to take Kyiw)
[1] https://www.startribune.com/for-tech-weary-midwest-farmers-4...
How big of a thermal battery (assume we are freezing water) would you need to store the cooling capacity of a 5 ton HVAC system running for 1 summer day?
Perhaps we could design a new generation of heat pumps with this approach in mind.
In Texas, we used to have residential plans where you payed wholesale rates for electricity (e.g. Griddy). I used to be on one of these plans and would constantly fantasize about being able to accumulate HVAC capacity at night when you would sometimes be paid to consume electricity.
(I do believe that switching to variable rates would be a good idea, that in even the poorest of the poor would the grand scheme of things be better off if they had to occasionally disconnect, less bad than if the alternative was the entire grid occasionally browning out, including services that might be more important to them than their home consumption. And peak prices would even be as high as they are now, if they weren't propped up by an army of fixed rate consumers who don't show any have of demand flexibility almost be definition)
Completely agree. Any time we try to control/subsidize costs we introduce instabilities and bad incentives into the marketplace.
Texas grid was about as close as you could get to reality for a while. I would much prefer a situation where everyone is impacted by the cost in the same way. At grid scale, nothing can be stored, so financial arbitrage is effectively a scam.
Run your A/C overnight when power is cheaper and the air conditioner is more efficient (and rolling blackouts are less likely...). Cool your house to say 4-8F cooler than you'd normally keep it (close off some vents in your bedroom if you need to, though personally I prefer sleeping in the cold). If your house is well insulated you may be able to make it through a significant portion of the next day without the air conditioner needing to cut in again.
In the grand scheme of things wood doesn't have a lot of thermal mass, but if there's a lot of it it still adds up. Even as the air begins to warm, the floor and walls still feel noticeably cooler.
Relevant docs I've checked for behaviour:
https://cloud.google.com/memorystore/docs/redis/high-availab...
https://cloud.google.com/sql/docs/mysql/high-availability
EDIT: Have found out from our ops team that the SQL instance recovered around 3am so it was down for approximately 9 hours -- which is still totally useless for something deemed HA.
From earlier in the incident history:
> Cloud SQL:
> Impact/Diagnosis: Non-HA instances backed by europe-west2-a are hard-down in europe-west2-a. HA instances that were in europe-west2-a when the incident started, are down with stuck failovers.
Try Spanner if one region is not enough.
Very frustrating when you're anxiously awaiting new information, and you have to do a word-by-word mental diff.
I think Google Cloud has one of the only status pages that is always up to date and very forthcoming in giving as much detail as possible. Personally I couldn't ask for more.
I think "you couldn't ask for more" is disingenuous at best, and an actively harmful outlook at worst.
The value proposition of "cloud" for is not "I can haz cloud of big corp as my own" but "I can haz many cloudz to make resilient infra!".
If you rely on one cloud provider you are doing it wrong.
It would have been rather useful had GCP linked these (presumably linked) incidents. The first mention of cooling in the concurrent incident was at 14:39 PDT (over 5 hours after first status update, and 4 hours after the cooling incident was created)... This is what was said:
> Description: A cooling related failure in one of our buildings that hosts zone europe-west2-a for region europe-west2 is affecting multiple Cloud services.
For instance, the idea that an expensive, thousand plus dollar rackmount server is only able to run in a special place where the temperature is just right and might fail if a fan fails or the temperature is a wee bit higher than usual is utter bollocks.
I build my own rackmount servers that can run at 100º temperatures, even with fan failures. I know this because that's how I test them. I have the OS aggressively throttle on the most egregious failures, but fan failure is much less common in general when you're using 80mm Noctua fans instead of 40mm fans that have to run at many thousands of RPM to keep their zones cool.
So maybe people need to rethink the idea that datacenters have to be kept at 70º or below, and instead should insist on better thought out hardware.
Yeah, servers can run just fine at 40℃. Well… unless they fail because of it. :-) That is: The ones that don't fail work just fine.
If you have a DC with 10'000 of your servers, and maybe 30'000 hard drives. What percentage of them will fail on any given day at 25℃ vs 40℃?
But it's not just your servers. Can your AC equipment/evaporators work at 40℃? And if your ACs start failing you could be looking at a cascading failure where it's actually more like 60-70℃, or just plain "a fire", in your DC.
Can your generators work?
And the answer also isn't "every component in my DC must be milspec extended temperature range". It's actually fine to build a DC in Iceland that's not specced for outside temperature of 50℃. In fact it would be a ridiculous waste to do so.
Google of course measures this. E.g. https://www.techrepublic.com/article/google-research-tempera...
But do keep in mind that 40℃ outside may mean 50℃ inside. Or indeed just 60℃ hotspots on the DC floor.
Hell, your network cables may not even be rated for 60℃. They usually aren't.
Your server may be fine with 40℃ inlet, but going out the air may disconnect it due to melting the network cable.
I bet Mumbai doesn't mandate winter tyres in winter, right? Sweden does. If suddenly one winter you see -10℃ in Mumbai, would you appreciate Swedes mocking you for not being able to drive in a car not designed for it, with tyres not designed for it, on roads not designed for it, etc…?
Nothing in the UK is designed for 40℃. Buildings, the type of steel made for train tracks, ventilation in tube tunnels, the asphalt, the walls in the building, the windows (no double glazing in Mumbai, I assume?). I would expect everything in Mumbai is designed to handle high temperatures. But not cold.
Company I work now in completely does not care at least based on me rising those concerns to mng.
Maybe that’s an obvious observation but I would have expected they had a little more operating range right at the point you need them.
Your fridge, rather than being the 5C it should be to keep your meat safe to eat, might have been up at 12C. You wouldn't be aware (it still feels cold), but you'd end up eating possibly dangerous food.
I really wish fridges had an alarm in that case (ie. The fridge has an indicator saying 'too hot. Food is now unsafe to eat').
There are also stickers for inside your fridge that can indicate the temperature. There is also a variant for specific temperatures, like 0, 5 or 7 degrees, that colorize if the temperature has risen, giving a (non-reversible) indication your fridge has been too warm.
Edit:
Did a quick search for you: https://www.tiptemp.com/Products/Rising-Time-Temperature-Ind...
EDIT: sibling comments about irreversible temperature monitors are also a great idea I hadn’t thought of! Time to buy some of those too
Heat waves indeed kill equipment. The system is attempting to deal with a higher heat load on the home/conditioned space, while also having to operate with high(er) ambient temperatures.
High ambient temps -> High heat load -> warm(er) return refrigerant temps/high(er) refrigerant pressures -> less cooling for compressor / higher mechanical load on compressor -> Elevated power consumption as compressor is working harder -> higher power/heat levels stress electrical insulation & components.
The system struggles to keep up as the duty cycle is elevated as well. Putting a sprinkler next to the condensor (outside) unit is a hack.
However, water companies have been telling people to reduce water usage during the heat wave since demand is higher than usual, so this may not be a great idea.
If so, This will only play into the hands of those setting up data centres in the far north of the northern hemisphere (e.g. Iceland).
"Oh, tanks by that foreign army are rolling into our cities - should we start thinking about a defense system?"
Global warming will affect every aspect not only of your business, but also of your life, maybe a little bit later when you are rich and can afford to live in a self-created bubble, but it will.
Yes, you should think about it. Hard. Now.
On-premise hardware still needs cooling, and is arguably harder to cool as there are fewer economies of scale on cooling infrastructure. Dedicated "bare-metal" machines are just in regular data centres so no difference to the cloud there.
I think data centre locations will still be chosen on two factors: distance to customers, and cost of energy. It's just that operators will be looking for cheap energy. Iceland is good because they have a lot of geothermal energy, not because it's cold.
It's also interesting that the status page's attempt to spin the scope of impact actually makes it seem worse that it was full-region outage (they said, "There is a cooling related failure in one of our buildings that hosts a portion of capacity for zone europe-west2-a for region europe-west2 ...")