Oh wow, I was a noogler there 4ish years ago, and they still tell that story as an example of "its okay to break things, we care about fixing them after so they dont break in the future"
There's also the one where all the frontend servers worldwide went into a crash loop from a bad configuration push. The SRE doing the push noticed some "weirdness" and rolled back even before the full scope of the issue was known. That one's in the SRE book.
0. https://landing.google.com/sre/interview/ben-treynor.html