Ugh, I wouldn't want to be that guy (even if there would be no direct repercussions). That said, and as others have highlighted - kudos on the writeup and openness.
Ugh, I wouldn't want to be that guy (even if there would be no direct repercussions). That said, and as others have highlighted - kudos on the writeup and openness.
1. Since we tested the change on a subset of A for a few weeks, we can assume it will work for all of A.
2. Since we tested the change on a subset of A for a few weeks, we can assume it will work for all of A and all of B.
#1 seems reasonable, but #2 is what needed to hold true in order for there to be no problems, since the change was actually enabled for all of A and B.
But was the engineer actually advocating to enable the change in B, or was that an accident during the manual deployment?
3. Since we tested this change on a subset of A and a subset of B, we can assume it will work for all of A and all of B.
That's the thing about these kinds of bugs...they are, by definition, tricky enough to have passed testing unseen.
These kinds of "almost took X offline" happen All The Time, its just that most of the time they get caught before it gets too far. Its inevitable that a few will squeak through the nets.
Mistakes can and will happen anywhere we allow them to. If you want to prevent mistakes, write tools to help reduce the "attack surface" (areas where mistakes can be made). Eg Don't want someone to be able to do "sudo reboot" accidentally? Alias reboot to something else. It won't stop hackers but it might help fight fat fingers.
Accidents happen.