It's seems these major infrastructure outages always are from a configuration change. I remember Google had a catastrophic outage a few years ago, and the postmortem said it all began as a benign configuration update that snowballed worldwide. In fact I tried googling for it and found the postmortem of a more recent outage, also from a configuration change.
Some seasoned sysadmin will say to me, "Of course it's always from a configuration change. What else could it be?" I don't know, it seems like there are other possible causes. But in today's superautomated infrastructures, maybe config files are the last soft spot.