Cold restart whole system after total outage
evalapply.org
evalapply.org
I have a friend who knows a lot about the phone system (he has a security clearance for some of his telephone work). One time we had a long conversation about this topic, until at one point he said "and let's talk about something else" -- I guess from that point some of the details are classified. So maybe there is a plan, or maybe they just designed the system in such a way that they could convince themselves that it would not go down unless things were so severe that loss of the phone system would not be your chief worry.
---
In September 2001 there was a full standdown of US airspace. That was accomplished pretty quickly: "you are ordered to land immediately on the closest airport that can handle your aircraft, or be shot down". Undoing that, however, took some careful planning! Fortunately the standdown lasted several days so there was time to work it out. Even if you had a plan for this (and I assume FAA had one), figuring out what the realities on the ground were and matching them up with the plan was nontrivial.
Apparently some of the planes landed where they could not take of again unless they were empty with a small amount of fuel to get to an airport designed for them. I don't believe I heard that any planes landed where they could never leave.
My belief from working in very large companies, and (previously) in mission critical systems is that a clean bootstrap and recovery process is extremely unlikely, almost impossible. Because in complex systems full of legacy parts and people who have long retired, the stars won't align.
The only way to truly know is to design and periodically test for disaster scenarios (emphasis on the plural). But due to the scale in time and space, cost and bureaucracy, this planning and rehearsing is not going to happen with the desired detail and intensity. People do not seriously plan for things that have never happened.
If it does happen, there will be a small group of extremely capable people that will find a way to bootstrap the system. It won't be according to some previously laid out plans -- they will make the plan in real time. They're not famous and probably never will be.
Well, more because the military industrial complex lines the pockets of your politicians, who in turn decide how to spend the budget.
https://www.rte.ie/archives/2018/0521/965058-mexican-lands-p...
In the merry month of May
Just before the dawn of day
A plane flew in for Shannon to refuel.
Because Shannon is fogged out,
Their are the rite of ought
To touch down in Cork Airport as a rule.
As he flew towards Mallow town
His supply of fuel was down
But the pilot was as cool as cool could be.
In a racetrack west of town
He made a safe touch down
Just beside the Mallow sugar factory.What a great story. Thank you for posting it.
The network stuff underpinning a lot of critical tdm phone traffic these days is like a collection of 23 year old Cisco 15454 held together by spare parts and a few people who care about them.
Yow. Way to make a guy feel old! :P
One of the weird challenges in building a new state-of-the-art inter city DWDM transport network now is dealing with things like legacy customers that have one OC48 and are unlikely to drop it any time soon, it's a considerable monthly revenue source, and have to deal with stuffing that into the system along with 100Gbps and greater coherent circuits.
Also from a customer relations perspective sometimes the customer literally forgets that they have this extremely exensive DS3 or OC48 or something in monthly recurring billing, and you don't want to bring it to the attention of management, because they might go "are we still using this?" and cancel it.
And break it does.
https://www.kron4.com/news/bay-area/911-dispatch-system-in-o...
It would have been better if it was really big SLA breach fines however.
Also, if that script remained valid at the time, I doubt it would do any critical actions. It might have been a sort of literal script to follow --- run the script, see what it says, do a thing, run the script, and so forth. Its supposed job was to help humans solve a bootstrap problem.
I see how my wording of that passage makes it sound like the be-all end-all of cold booting a telco. But that's what we get when we wall-of-text in our Slacks :D
(edit: clarifying remarks)
This has happened before outside of 2001, albeit not really a DR issue but more a political issue -- if we look to the Meigs Field destruction by former Chicago mayor Richard Daley, multiple aircraft were left stranded with a now destroyed runway. (The solution was to just give them special clearance to take off on a taxiway, but still.)
https://web.archive.org/web/20110720045652/http://www.aopa.o...
For some reason, he decides to demo the UPS cut-over switch. I have no idea why. But he manages to toggle it the wrong way and instead of switching the entire room full of servers to the UPS, he manages to cut power to _all of them_.
My recollection is that the cooling went out too and the room was suddenly very silent. But in retrospect that doesn't make sense.
What I do remember is that it was non-trivial to bring all our Unix servers back up because over the years they had been setup with NFS mounts in a loop such that for A to boot, it needed B to be up, which needed C to be up, which needed A to be up.
Oops.
So it took a lot of manual intervention to bring everything back up.
He was really nervous showing the new install off to us, and he was just talking through the positions of the switch; but he actually moved the switch to each of them as he did it.
Back in the days the conslops ran the asylum. :)
Any old school UF types looking at this, The Dog list lives on, and is again meeting from time to time. :) But at quieter venues. Our ears have unaccountably gotten old.
We were student employees, who were nonetheless in charge of all the systems. It was an absolutely awesome experience.
Sorry, couldn't hold myself :)
Cisco had a virtual router. N1000 perhaps, the hardware was a C220 and it had some sort of appliance running on it. Which depended on some sort of LUN from NetApp for its storage. But you couldn’t stand the LUN up until the vCenter was up because UCS provisioned the LUN, it didn’t work if you did it manually, and that ran on vCenter, and the vCenter depended on the router so that it could reach the LUN. It was pure circular logic hell.
There was a path to make this work, but it literally took half a dozen very ~~clever~~ expensive Cisco and NetApp engineers a couple of days in a room with a whiteboard to figure it out. It was absurd.
Second story, around 2005---I, along with a friend and my father, were in Las Vegas eating lunch at one of the major casinos when the power went out completely. It was eerily silent and dark! (and then slowly, we started hearing the groans of slotzombies rising among us) I'm sure someone lost their job for a UPS cut-over failure.
It seems like if there really was a disaster, first of all nobody would know that script existed, and second of all if they tried to run the script, it would fail because of all the changes to the system since the script was initially developed.
Isn't there some saying like "if you don't test your backups, you don't have backups" or something like that?
If the entire telephone system was down and needed cold started, and the script had information someone needed to do that, someone would take the time to read it. Maybe not run it, but definitely read it to extract clues.
I mean, it's not like it's binary. It's totally possible.
Can you even imagine the alternative? "Hey lets throw this maybe incredibly helpful shell script in the trash. Because it's too long."
it’s just utter bullshit.
In a disaster scenario, something is better than nothing
Also of course it will be crazy to read a giant shell script. But then again if the stakes are high enough, and if it yields even one critical piece of information, then it's worth it.
The larger point is that organisational knowledge clings on in strange ways. In a crazy disaster scenario, people may appreciate having access to anything they can get their hands on.
Classic problem of deferred costs. Backups cost money and it's tempting to avoid investment in them (ie fail to test them) but that can bite you when its least convenient.
I say excellent, not because I assume it would have been well written, but because when it comes to these things I prefer code (any form of code, indeed) that I can be reasonably sure worked at some time over natural language documentation any time.
I completely agree with Dr. House when he says he never asks the patient, because they lie all the time. Same with human written documentation.
> So much of the modern world depends on our mastery over materials (to make a precision screw, you need a precision-machined harder material—diamond / titanium—to work on a softer material—steel), and our ability to turn rotary motion to linear motion (it's stupidly difficult to reliably precision-machine a harder material without even more precise linear + rotary motion—lathe/CNC machine). Hence, a bootstrap problem.
Steel is hardenable (or rather, some steels are hardenable), you can change its hardness through the specific application of heating and cooling. So you can make a crude tool with relatively soft steel, harden it, and use it to make a more precise steel tool (again machine soft, then harden). This does make the bootstrapping problem a bit easier, I think. Although not easy in the absolute.
See https://www.youtube.com/watch?v=V_Mp1fNzIT8 for a great dive into primitive steel hardening techniques.
And there's way to make three perfectly flat sharpening stones by starting with three raw pieces of natural sharpening stone, just by alternately rubbing the three stones together until they flatten each other out.
Paul Sellers can teach you how to flatten a large board without a planer. He also has videos on how to get a wood plane perfectly flat using a large sharpening stone (which can be made as above or with float glass).
And if memory serves, you to make something perfectly round you first need something perfectly flat. Once you have something perfectly flat and something perfectly round it's off to the races.
Edit: "The Origins of Precision" is a half hour well spent https://www.youtube.com/watch?v=gNRnrn5DE58
Flat boards require a flat plane, and like chisels, the tolerances on a new plane are fairly loose. Partly down to thermal contraction (from running the production line too fast? I've never gotten a straight answer). So the first thing you do with both is grind them truly flat, and you need a reference surface for that, like float glass or a diamond stone. Common protocol is to use the diamond stone only to flatten sharpening stones, and the sharpening stones to flatten chisels, and the chisels to flatten mortises. Basically diamond stones are very accurate but too expensive to have sufficient grit ratings and longevity
This is not the one I'm thinking of, but it's a taste:
https://www.youtube.com/watch?v=Cl5Srx-Ru_U
Right angles can be achieved by a process of iterative refinement. A square is two flat surfaces that are used to adjust two other flat surfaces, and they are only at right angles when 90.0º + 90.0º = 180.0º. So if you reverse the square or make two identical squares, they should touch along their entire length. If they don't then they're not square. Alternatively you can apply a square multiple times and check if the 1st and 3rd plane are perfectly parallel. Or if the 1st and 4th plane intersect at the same point, which also increases your accuracy by 4x by multiplying the error. I've seen this demoed by fine woodworkers squaring up a table saw for instance.
Going from iron to precision screws is a matter of first making precision flat surfaces, then lathes, and onward from there.[2] You can do that with just iron and heat treating, but it won't be easy.
If you want an alternate history where something slightly less drastic is dealt with, the book "Ring of Fire - 1632" by the late Eric Flint[3] is an interesting place to start. In the book, a town from West Virginia circa 2000 is thrown back into the middle of the 30 years war in Germany. Lots of exposition of the book is about the supply chains we all depend on, and how they work. It's the start of an awesome series.
Books and working knowledge, are a precious resource. As long as we have a critical mass of them, and conditions remain reasonably tolerable for human life, we can recover.
[1] https://www.youtube.com/channel/UCAL3JXZSzSm8AlZyD3nQdBA
[2] https://ia800104.us.archive.org/20/items/FoundationsOfMechan...
https://en.wikipedia.org/wiki/List_of_books_in_the_1632_seri...
I quit the storage / backup industry because 7:30am phone calls would make me hyperventilate in a panic.
> I have seen this at <Indian eCommerce Giant> and at <a FAANG>. Most of it is related to cached data. Cold starts with empty caches causes too much load on databases. And then the failures cascade.
> — Another M'colleague in the Slackroom.
Isn't that not really a problem with cold restart per se, but more the restart procedure? If caches are so critical, wouldn't you need a feature to throttle the load to what the databases can handle, as the caches populate? E.g., if you're cold-rebooting Facebook, start by blocking all connections except those geolocated to North Dakota, then add other regions as your caches fill.
[1]: https://www.youtube.com/watch?v=30jNsCVLpAE
[2]: https://www.tritondatacenter.com/blog/postmortem-for-outage-...
So then question would be how to bring back free-floating currency, if we somehow shot ourselves back to the dark ages and forgot all about it.
"Lots of luck" would be my best guess. (Because if we manage to go back there, we are likely to succeed at going back further and forgetting even more. These sorts of swings have multi-generational momentum I suppose.)
> The author appears to have missed cyclic dependencies as a barrier to cold restarts.
P.S. I am sure I have missed plenty in that stream-of-consciousness post (and mistaken a bunch too)!
There was a story around this during the iraq war where a US military virtual machine system when down, and had to come back up without internet. Problem was VMware needed DNS to start the VMs, one of the VMs it needed was Active Directory for security, AD hosts the DNS and now you're locked up without an external running system.
DNS itself is typically a cold start nightmare.
My earliest exposure to this concept as a child was watching the film Jurassic Park. As someone fascinated by systems I found the idea of having to bring the whole system back up from scratch pretty interesting.
Today I still find these kinds of bootstrapping processes fascinating - both these megascale processes, but also the boot process that occurs whenever you turn on your computer. The latter is probably one of the most Rube Goldbergian feats of engineering with us today that actually still achieves a useful purpose. In fact it's absurd how Rube Goldbergian it is. And the complexity of the boot processes for modern systems (see [1] for a small glimpse) is extraordinary.
When you turn on your computer, it's like you're re-executing the entire process of a civilization bringing itself into being, gradually developing progressively more sophisticated technologies: at first RAM isn't working, but then you get RAM working and that lets you get progressively more sophisticated parts of the hardware working, etc.
By comparison, humans have no "automatic boot process". We're constructed in the 'on' state via fork(). So this repetition of entire process of, ah, 'abiogenesis' whenever you turn on your computer is kind of insane by comparison. Entire kingdoms of hardware state rise and fall with the press of a power button.
As an aside, I'm fond of the Red Dwarf novels, which are set on a massive mothership-type spaceship. The ship was constructed in space and never designed to enter a planet's atmosphere. In particular, the ship, and its engines, was constructed in the 'on' state by the crew that built it originally. It was never conceived that the ship would ever need to be rebooted, because it was assumed once the engines were initially fired during the commissioning of the ship, they would never be turned off until decommissioning. Thus, the ship has no automatic boot process for the engines, only a manual engine firing procedure which is extraordinarily arduous and long-winded and takes weeks to execute, said procedure having been included in the manual only as a curiosity more than anything else. This idea of a ship built "on" under the assumption it would never once be shut down or "restarted" until decommissioning is interesting, but of course also directly mirrors biological life.
"Push it."
When Spielberg was on he was ON! Who else could take a scene like "they have to reset the circuit breakers" and make you absolutely on the edge of your seat over it.
It turns out a lot of high-power circuit breakers need the energy from a clockwork mechanism in order to open in the event of a fault. So the thing where you have to pump the handle is actually to charge a clockwork mechanism to ensure there's enough mechanical energy to open in the event of a fault - AIUI. Presumably, the breaker is designed with safety in mind and won't let you push-to-close until you've done this.
I always thought that the main breaker looked much more real than the "individual park systems". Turns out it was! The other "breakers" with the backlit names are definitely from the prop department
Maybe some worms are, but "we" most certainly aren't. The closest analogy might be execve("/proc/self/exe"), but even that is flawed.
Likewise knowledge around nitrogen fixation and fertilizers.
There are probably a half a dozen huge improvements that could be made for bootstrap_society_v2.sh
Perhaps we should all write down our version and tuck it away somewhere safe just in case. Maybe on something more durable than paper. And certainly more durable than electronic storage.
An amusing aside is that IP telephony sometimes gets into a mess here, as SIP sends a 503 with a retry after prompt to the client, but its not always randomized, so you can get these waves of barbarians at the gate, who all go away, and then all come back again exactly 180 seconds later...
Then again we couldn't cold start a supply chain or a semi fab or humanity itself so maybe that's the default.
Similar premise, but focusing more on philosophy and understanding, and not tracking linearly through time.
There were another two Connections mini-series, though they were IMO not up to the mark of the first. Good, but not truly excellent.
It doesn't just speak well for disaster recovery prospects (both the feasibility of doing it and the density of developers who could possibly pull such a thing off), it's also very, very useful for speculative development.
When you make a high barrier to entry of making large modifications to the system, you also tend to create an underclass of developers, who never really get to understand how the system works.
What if we split these two microservices into three, or combined these three into two? That's a pretty common question, that only gets asked if you know you won't get laughed out of the room for suggesting it.
A history of the screw. Really interesting around how it was developed. Some machining techniques are far older than you’d expect, and some capabilities far newer.
The books thesis is that the screw is the most important invention.
> Even though nothing will go as planned, it's important to have the memory and expertise that did the planning, because that's what's going to be able to think through the as-yet- unknown-unknowns, when the inevitable FUBAR situation suddenly happens later.
We don't plan so that everything will go according to plan, we plan so that we are better equipped to reason when the plan doesn't work.
https://practical.engineering/blog/2022/12/5/what-is-a-black...
How could one possibly balance the load with the plants coming online? If the generation and load is too mismatched, the generators can literally automatically trip off the grid, so generation and load must be carefully balanced as things get brought back up.
One would almost need to shed nearly all the loads from the black grid (Which may have happened anyway as the grid collapsed, but any loads not already shed by the collapse could prove interesting), and re-add some some gradually as plants come online, which still seems crazy difficult.
And inrush current demands from many loads as they are get reconnected must be pretty insane.
Something about bringing up a designated plant and feeding the output over 'cranking' lines that other plants along the route can synchronize their output against. Then gradually adding load and source until the system is meta-stable again.
Edit: additional data
Not only synchronize, but also use for internal needs like all the pumps and particularly the 'excitation current' that establishes the magnetic field for the generator. It allows control over the output voltage. There are also other drawbacks to the more obvious solution of fixed magnets which can be oversimplified as 'ware'.
https://en.wikipedia.org/wiki/Permanent_magnet_synchronous_g...
The trickiness seems worst close to the very beginning when even relatively small misestimation of a chunk of load being restored would have a proportionally bigger impact. Many loads are not completely predictable, so presumably they would need to favor bringing some of the more stable loads online early so that normal variation from the loads that can only be predicted well in aggregate won't vary enough to trip everything back offline.