The process is "panicking", but it is still reliable enough to post a message to Slack? Even writing a log entry might be too much, depending on what exactly the failure was.
The process is "panicking", but it is still reliable enough to post a message to Slack? Even writing a log entry might be too much, depending on what exactly the failure was.
Well.. most likely yes. If I invoke panic, all state is exactly as it was just prior to that statement. And even if not, it's not gonna "launch the missiles" if that slack post or log-write fails now, is it?
> Even writing a log entry might be too much
"Too much", how? Either it succeeds in which case it'll help you, or it won't which has the same result as not attempting it in the first place: no log entry..
Meanwhile, the crazy code that panicked is still running it's other goroutines - remember, you called recover! - so maybe now the webserver still has an open port and is allowing your users to access whatever strange state is left inside... gonzo things really can happen if you let a corrupted program stay on rather than shutting it down immediately. Even if nothing bad is happening, you still are out of commission for that entire period.
The point is, murphy's law always comes into play. So if we're talking about production best practices, consider that "most likely fine" means "definitely not fine at scale over time". Just make sure that whatever you're doing during shutdown can't block.
I'm seriously wondering if you ran into any trouble with recovering panics in production, because that would imply all Java, C#, Python, JS and Ruby server code in production which is happily catching and logging exceptions in the main request handler is constantly running into corrupted state.
At every job I've worked at it was generally discouraged to handle MemoryError exceptions in python and OutOfMemoryError exceptions in Java. It's safer to assume there is nothing you can do on such occasions, because sometimes (usually the worst possible time) there isn't.