> This is a lot like software developers and their persistence. When your production service has a mysterious error in the logs, do you shrug and say “well, probably some transient network bug, it won’t matter” or do you dig in to understand what’s going on? Both approaches are reasonable, and it takes a balance of pragmatism and curiosity to be effective.
I find that it really depends on how harried I am. Am I (or someone on my team) able to dedicate time to this issue, or do we have other higher priority concerns?
Every single time I can think of where we've had a "that's weird", it was a bug in the software. But it would take days to weeks for someone find the root cause (these were 1 in 100,000 race conditions), and we just can't afford to dedicate that amount of time to each.
So we have to pick our battles.