About the closest I came was when working in telecom, for a new tech deployment I had my hands pretty deep in the lab environment, so when things were broken in the lab, bad config, etc I would get pulled in troubleshoot and solve the problems.
Well one day 911 wasn't working in the lab and the problem got thrown my way, and it wasn't an obvious problem like someone miss configured something or broke some config somewhere. In telco at the time it was all vendor driven solutions, so I intentionally left the system broken to bring in the vendor to troubleshoot, and it was clear, this is a lab, let's not treat it like production and as such we don't need immediate recovery, we want to get to the root cause so it doesn't happen in production.
The next day, the handset verification team was on me, saying they need this to work immediately since they need to validate some device by such and such date. And I basically said listen, there's a software problem in this product, and we don't want it to go to production. And if I don't get it fixed it could blow up in production on us. I also told them if I don't make progress in a day or two, I would try and reconfigure another environment for them so they would get unblocked, but otherwise was not willing to just reset this system so the problem went away.
I was also doing my own investigation as much as I could since the vendor wasn't always the most reliable, and I encountered something unexpected. It looked like a node was rebooted, so I tracked that down, and found a senior architect who new I was working on solving the issue had rebooted one of the blades. His answer was basically the device team was complaining so he just went in and rebooted the node so they would stop complaining to him.
Luckily, he didn't know enough on how to really reboot the system, so it just synced back with it's backup and still had the problem for us to investigate.
The vendor comes back and goes ah yea, the 911 handler is using the wrong memory region for storing emergency calls, so instead of being able to allocate a hundred thousand records or whatever it was for active emergency calls it was using an administrative region that could only allocate something like 5 calls. This was enough years ago that I forget the exact number, but it was less than 10. Not just that, but there was a second bug, a certain 911 call flow would allocate the call but not release it, which is why we couldn't make any 911 calls in the lab, we had leaked all of the reserved memory for emergency calls.
I just remember being so livid, because the culture for anyone who dealt with that system was it's failure is just in their way, so lets just escalate and try and make it go away so we can continue on with our jobs.
And it would've been so easy to just reset the whole thing so that people would stop complaining. It was just a lab after all.