276 karma · joined June 27, 2019
True or false, the jokes write themselves sometimes
What is this “life” you speak of?
> That's not hard to debug
I agree there's a large class of problems where that's the case, and I also agree it's perhaps easier than many of the issues we've caught in the real world, such as are discussed in the Cisco case study. But I still want to defend the example as non-trivial.
So of course we pick up the errors in services that are fall-out from the breakage, and we pick up the chaos test running, which leaves a pretty wide footprint. But if something like this happened in the real world, the important event would have been that one that's in the RC report, that's not even an error (though it is quite unusual): the kernel log entry pointing out the eth0 misconfig. It can take a long time for someone to get around to poring through their various host kernel logs and looking at everything, even non-errors, so having it surfaced to you right away feels like a very useful thing.
At the same time though, I like the example you gave even more. What many people will do is upload their own incident data via the CLI, or deploy our chart into a staging environment and just break things, to see what happens. Based on your description, there's a good chance we would pick it up since we track anomalies and perform anomaly correlation across all pairs of streams: we're not looking for a threshold percentage of overall error rates, for example, anywhere. If you'd be willing to help us test this use-case, I'd love to work with you personally to help you get up-and-running.
Also, as we see more and more failure modes, we continue to make our detection algorithms more robust. While we generally achieve >>90% detection of incidents and their root cause indicators overall, there are always places we can and would love to do better. I think you'd find that our solution would catch most of what you'd want out-of-the-box, and that we're responsive enough to learn from every customer.
Lastly, I would ask: do you think that we should do videos of some more "subtle" examples (for lack of a better word)? As someone viewing the website, would you have watched them?
Thanks again for your feedback!
First: in both cases things got better sooner than I had expected; within a few years things were good... once I had dusted myself off and applied myself to the next chapter of life, I was able to take steps, find ways to dig out, and find hope and passion again in my pursuits.
Second: I, like you, had a lot to offer the world, my profession, and my family and friends; I look back and can see these things. Please, hang in there, and please, don’t let the downturns extinguish your light. Whether you know it or not now, many people will need you around, and in the end, even more will be better off for your influence.