Our team managed to screw-up some pretty major DNS due to a valid terraform plan that looked OK, but in reality then deleted a bunch of records, before failing (for some reason I can't remember) before it could create new ones.
And of course, we forgot that although we had shortened TTL on our records, the TTL on the parent records that I think get hit when no records are found were much longer, so we had a real bad afternoon. :)
lifecycle {
create_before_destroy = true
}
may be your friend :) (not sure if applicable though)See a similar outage in S3 from 2 years ago - https://aws.amazon.com/message/41926/
That above seems pretty clunky, so it's very likely not what happens.
In this particular case, commands that were run on a Production machine were by-design limited to what they can do and affect (mostly just the physical host they’re run on or a few hosts in the logical group of hosts they belong to).