Outage Stories: The copy and paste outage
wdkwwdk.com
wdkwwdk.com
STONITH is a good design option, but I think this implementation would've still struggled. The way they were tracking processes was fragile, so the shoot the process lost track of which pid it should signal. And this was a sort of clustered system, so if you shoot the node you take out other services (think of murdering a kubernetes node to terminate one process, you're going to inflict some spillover.)
Finally after a few minutes,our head of IT/networking/everything finds the tech's personal cell phone, and calls him from his own personal cell phone, because VOIP is dead in the water now. Tech unplugs network device from network, within 5 min everything self-heals. Untold amount of fines and fees for not being able to handle trades during the outage.
1. Configuring VLANs is hard, let's turn on the Cisco feature to sync VLANs 2. Network switch from lab or another network would be moved to production, with a higher revision index. All switches sync from the highest revision. 3. Outage happens because VLANs in lab are not same as production. 4. Investigation says vlan trunking protocol is dangerous, and should be disabled. 5. A year later, forget about the outage and repeat.
I think eventually there was a feature that sort of solved this problem, but I don't remember after all this time.
> The supervisor script just starts detecting the old process, and has no idea that the additional instance has been launched
The number of times I've seen outages and weird flaky problems from some shitty script trying (and failing) to parse ps output...
daemontools, runit, s6, systemd, etc all solve this problem properly.
Just feels dirty or something.