We unplugged a data center to test our disaster readiness
dropbox.tech
dropbox.tech
The library was created in .net framework 2.0 and no longer exists or is maintained. It’s been on the list for 8 years to replace…
It has a memory leak which causes IIS to consume gigs of data over time, resulting in IIS completely restarting itself and causing all hosted sites to go offline for 30s. so I set the IIS pool to recycle every hour and it hasn’t caused any issues in ~ 5 years. Due to this I feel like it’s dropped even lower in priority. :(
https://www.theregister.com/2020/04/02/boeing_787_power_cycl...
How many times have we rebooted our machines this week to fix some weird issue? (or maybe my OS install is buggy, but still :P)
$ uptime
19:14 up 42 days, 6:49, 11 users, load averages: 7.85 8.24 8.07
Zero times this week (or month).
(And since everybody is now going to share war stories... When I worked at Con-way, now XPO Logistics, they had the pop-tart incident, where a kitchen mishap triggered fire regulations that mandated the data center be taken down, and since they had absolutely *zero* routine in restarting their data center, folks got to debug the backup plans in real-time, until two weeks later all service had been restored. Kind of. Me suggesting we should test this more often was only perceived as a joke in bad taste. I thought it was a swell idea, since our Sundays were scheduled downtime anyway. Yep, the whole day. Every one or two weeks. Suffice to say, it wasn't a very long stint.)
It's like accidents during fire drills, they happen, yet it's worth doing all things considered.
These exercises happen several times a year.
These days the team running them announces that it’s happening in an opt-in announcement group at 8am, pulls the plug at 9am, and barely anyone even notices because the automation handles it so gracefully.
Mostly I just miss the t-shirts, as the <datacenter>-storm events got the coolest graphical designs...
The first time we did it was painful. It took all day, probably around 40 people involved total. We found all sorts of problems, and we actually only failed over the production site, nothing else.
We'll be doing it for I think the fifth time in a few months. Mostly automated now, last year was pretty smooth, aside from a couple bad assumptions that crept back in to some code.
It is like anything else, repetition makes perfect.
> The requested URL /infrastructure/disaster-readiness-test-failover-blackhole-sjc was not found on this server.
> Additionally, a 404 Not Found error was encountered while trying to use an ErrorDocument to handle the request.
I see which data center was hosting the article, then.
Internet Archive to the rescue: https://web.archive.org/web/20220426191128/https://dropbox.t...
I might run three deployments at each data center: the primary, and secondaries for two other regions. Replicate between them at the block device level, bypassing the mysql replication situation entirely (except for on-disk consistency requirements of course).
Of course this comes with a 3x increase in service infrastructure costs because of the two backups in each data center that are idle waiting for load.
https://dropbox.tech/infrastructure/atlas--our-journey-from-...
Welcome to basically any large-scale enterprise.
I have grown to learn that the active-passive strategy is the best option if the business can tolerate the necessary human delays. You get to side-step so much complicated bullshit with this path.
Trying to force active-active instant failover magic is sometimes not possible, profitable or even desirable. I can come up with a few scenarios where I would absolutely insist that a human go and check a few control points before letting another exact copy of the same system start up on its own, even if it would be possible in theory to have automatic failover work reliably 99.999% of the time.
If your failover is happening "all the time" you basically just have a single system with failures.
> However, because of that choice, replication between regions is asynchronous—meaning the remote replicas are always some number of transactions behind the primary region.
Switching the active is a data loss event.
Or, you can choose to drain the load and wait for the replication to catch up, but that's downtime and not a test of the real failure mechanism.
Active-active is also valid, but the point is that it comes with a huge amount of increased complexity. At some point you need to make a value calculation to decide if you want to focus on that, or on building the product.
>Given this context, we structured our RTO—and more broadly, our disaster readiness plans—around imminent failures where our primary region is still up, but may not be for long.
IMO if the big one hits San Andreas, the SJC facilities will likely go down with ~0 warning. Certainly not enough time to drain user traffic.
It's interesting to note that Dropbox realistically can probably tolerate a loss of a few second to minutes of user data in a major earthquake, but cannot tolerate the same losses to perform realistic tests (just yank the cable, no warning).
If the earthquake hits at 3am in SF, it'll likely take both the metro and a significant number of the DR team out of the picture for at least a period of time. Surviving that kind of blow in the short term with 0 downtime is a very hard goal.
Is that correct? That seems very single-point-of-failureish to me.
Later in the article they do say 'began to reconnect the network fiber', which implies multiple connections to me. They also took reference photos and ordered backup hardware. So I definitely don't think it was a single connection.
https://onezero.medium.com/survival-of-the-richest-9ef6cddd0...
unpaywalled: https://archive.ph/AABsP