How would people react? What would engineers do to recover? I always found that idea fascinating.
Imagine Google saying tomorrow that they lost all the accounts and emails. What kind of impact the world will have?
How would people react? What would engineers do to recover? I always found that idea fascinating.
Imagine Google saying tomorrow that they lost all the accounts and emails. What kind of impact the world will have?
You not only have backups in place, you have documentation in place, including a back-up vendor who has copies of the documentation and can staff up workers to get it up and running again without any help from existing staff.
And we tested those scenarios. I'm not sure which dry runs were less fun - when you got paged at 3 AM to go to the DR site and restore the entire infrastructure from scratch... or when you got paged at 3 AM and were instructed to stay home and not communicate with anyone for 24 hours to prove it can be done with out you. (OK, so staying home was definitely more fun, but disturbing.)
How many companies really plan for an event where their entire infrastructure goes offline and their entire team gets killed? Does even companies like Google plan for this kind of event?
But based on my experience, the initial recovery planning is the hard part. The documentation to tell a new team how to do it isn't so painful once the base plan exists, although you do need to think ahead to make sure somebody at your back-up vendor has an account with enough access to set up all the other accounts that will need to be created, including authorization to spend money to make it happen.
In some ways losing your ERP and it's backups would be harder to recover from than both sites burning down, insurance would cover that at least.
It's hearsay, but I was once told that achieving "black start" capability was a program that took many years and about a billion dollars. But they (probably) have it now.
It's most often referred to in the electricity sector, where bringing power up after a major regional blackout (think 2003 NE blackout) is extremely nontrivial, since the normal steps to turn on a power plant usually requires power: for example, operating valves in a hydro plant or blowers in a coal/gas/oil plant, synchronizing your generation with grid frequency, having something to consume the power; even operating the relays and circuit breakers to connect to the grid may require grid power.
The idea here is presumably that Google services have so many mutual dependencies that if everything were to go down, restarting would be nontrivial because every service would be blocked on starting up due to some other service not being available.
Since 9/11, more than you might think. For example Empire Blue Cross Blue Shield [1] had its HQ in the WTC.
https://www.computerworld.com/article/2585046/empire-blue-[1... cross-it-group-undaunted-by-wtc-attack--anthrax-scare.html
And what a blast from the past:
> Some of the temporary locations, such as the W Hotel, required significant upgrades to their network infrastructure, Klepper said. "We're running a Gigabit Ethernet now here in the W Hotel,'' Klepper said, with a network connected to four T1 (1.54M bit/sec) circuits. That network supports the code development for a Web-based interface to the company's systems, which Klepper called "critical" to Empire's efforts to serve its customers. Despite the lost time and the lost code in the collapse of the World Trade Center towers, Klepper said, "we're going to get this done by the end of the year."
> Shevin Conway, Empire's chief technology officer, said that while the company lost about "10 days' worth" of source code, the entire object-oriented executable code survived, as it had been electronically transferred to the Staten Island data center.
It's pretty much part of the basic day-to-day life in some industries.
I sat in on a DR test where the moment one of the Auckland based ops team tried asking the Wellington lead, the boss stepped in and said "Wellington has been levelled by an earthquake. Everyone is dead or trying to get back to their family. They will not be helping you during the exercise."
https://www.nytimes.com/2014/11/19/magazine/the-secret-life-...
They had backups and were able to recover data and systems.
By the time I got there, they were somewhat functional.
The biggest problems were the lack of knowledgeable personnel, not lost data or systems.
I was there as a consultant and didn't know anyone there when I went.
I won't provide any details out of respect for those fine people, but the grief was so thick, you could have cut it with a knife. As I said, I didn't know anyone who was there (or wasn't there) but after a day, I wanted to cry.
HN comments at the time: https://news.ycombinator.com/item?id=487497
The site relaunched a month later and shut down for good a year after that
Slack is mostly real time communication, at least for me. There are a few bits and bobs that really should be documented that are in the messages though.
The rest of the world may not be so energetic re: their accounts and data, so it would be painful for many, it depends on their how much risk they are willing to experience.
Being in DR, it is very difficult for businesses to allocate the time and resources to good planning - for many, DR is an insurance policy. Staff: engineering and development are focused on putting out fires however, a real DR is more than most companies can handle if they have not planned accordingly or practiced through testing failover/normalization processes as well as performing component-level testing.