Is it just a (non)-survivorship bias where we most commonly talk about the failure cases and they actually have a stellar 99.99999% record that doesn't make headlines?
Is it just a (non)-survivorship bias where we most commonly talk about the failure cases and they actually have a stellar 99.99999% record that doesn't make headlines?
My best story was as a young student turning up for my helpdesk shift at about 5:50am to a phalanx of fire engines and the Hazmat team.
The generator in the basement was maintained and tested. This time it had started when power went out, but there was a pump that filled a holding tank from a 10,000ish litre tank under the building (speced to run for a week apparently). A pipe had cracked, and so it just kept pumping fuel on to the floor, carpark and into the local stream ...
The only person in the building overnight was the operator (tape changer and report runner) some 5 stories above all this, who eventually smelt it and eventually investgated. It was a massive nightmare, and the place stank for weeks.
Then there was the time the halon system went off unnecessarily...
This is what anti-siphon valves are used to prevent: they are valves that open only when sufficient suction is present in the fuel line leaving the tank. Thus, when the line breaks or the pump fails, the valve will close.
> or it can be a few floors about the tank, too high to suck fuel from the tank so it has to be pushed up from below.
True, it has to be pushed up then. Alas, you can have a sequence of suction pumps. Note that otherwise you'd need to have high-pressure fuel piping, which might cause more problems than the possibility of feeding a fire indefinitely.
Switching paket based networks is often much easier because you can just hold on to the packet and retransmit if a line fails, that's why it's much rarer to have total network failures.
Generators themselves are very reliable machines. The engines are cast iron (instead of cast aluminium as found in all cars since two decades or so) and are nowhere near as close to the edge of material capacity compared to car engines. Electrical generators are essentially just a big blob of metal, they are simple, efficient and very reliable technology giving many decades of service. Generator controllers are designed to be reliable and aren't terribly complicated, either.
If generators really are just a giant blob of metal and very simple (and I have no reason to disbelieve the GP comment)... well... it could be kind of interesting to build an open source software stack to handle switchover. Because, disclaimers notwithstanding, the code would ostensibly be super simple too. So even if it couldn't officially be used directly, it would certainly provide a good base for engineers to copy-paste, either literally or ideologically (and then thoroughly verify, of course!).
Okay... thinking about it, I'm probably wrong - either generator control has some fundamental intricacies that make it not-completely-simple, or all sites have edge cases that have to be baked into the firmware by a system integrator/electrical engineer.
I say this because I'm (genuinely) trying to figure out why the PLCs failed - both in OVH's case, and in AWS's case (see elsewhere in this thread, https://news.ycombinator.com/item?id=15676189).
It's obvious there are crazy but legit reasons for this kind of thing to happen. I'm very curious what the complexity scale here is.
I've seen what switching 20kV looks like in videos - yeah, that kind of thing requires very careful design, and is invariably going to come entangled with a PLC-style controller, as is the norm for industrial equipment.
Disaster backups are hard to test as to do it properly they have to be integration tested, and doing so might break production. Sure your generator might power on for 5 minutes a month and make people happy, but there are many unforeseen things that are hard to test for.
In our case, we had a generator run out of fuel. Another time we had one have the battery die (they're started by batteries just like a car engine)
Every day. All of them (automated).
It would be interesting to run a datacenter completely on batteries that are charged from solar + the grid + gas generators. You could run the generators at optimal power output and use them daily. The batteries required for ~24-48 hours might still be too expensive to make this possible. Maybe some kind of super low power datacenter could pull this off. One day it shall be mine.
Like others said, yes they are big and with lots of moving parts. But if you take care of them (every month by the book) no one should have any unexpected problems.
For me is mind blowing that OVH did the last power failure test in MAY. WHY so rare? IMO that is 1st mistake.
They do seem to test the generators every month.