Anybody want to guess root cause?
Do we have a "root cause" bingo card?
DNS
Database
What else are super likely?
Anybody want to guess root cause?
Do we have a "root cause" bingo card?
DNS
Database
What else are super likely?
I often have to work hard to convince people of all experience levels that it's the best way forward.
- "It's just a little bug I can just fix it [and definitely won't make it worse with code that I haven't tested as rigorously right?]"
- "My KPI/bonus/project plan relies on this going out today"
- "My code is fine it's the infrastructure [that I didn't warn] that can't handle it. They need to fix their side now."
I don't know about your VP but "how fast can we get back to before it was broken?" is reasonably the first thing you should be asking
Doing #1 puts some serious boundaries on how bad it can get
Before finding out the dead simple failure mode and fix, engineers need to spend countless hours diving into the most technically complex scenarios that might be happening but are irrelevant. Then they can reset permissions or add disk space or add back a DNS entry.
Won't save you if someone's running-as-root reporting job goes rogue and fills up the disk, though, while the file might... I mean, obviously one ought not have done that in the first place, but the real world is a whole thing.
Busted API deployment?