Show HN: Using machine learning to recommend dashboards during incidents
beta.overseerlabs.io
beta.overseerlabs.io
When we first started working on this project over a year ago, we weren't sure if the algorithms would work, or if our insights would be of value to anyone. We were also struggling to figure out how to make it easier for people to try the product without having to change their existing workflow.
Since then, we've made huge improvements to the algorithms, deployed the tech for several large customers, and demonstrated value. Now I'd love to get a bit more feedback from you guys and see if we're going in the right direction!
So here's how the tool works: 1 - We pull down your dashboards from your existing monitoring tool (e.g. Datadog/Wavefront/Librato) using your API key. 2 - We integrate with your PagerDuty account via a Webhook to notify us when an incident has triggered. 3 - When our Webhook is invoked, we will use machine learning to rank your dashboards, rank the metrics on those dashboards, and notify you via Slack/Email of the top dashboards/top metrics on those dashboards to look at.
For this demo, we only expose the Wavefront plugin, and you'll be able to configure it on the initial page.
To integrate with PagerDuty, you'll need a URL to our end point, and we'll need an email address where we can send the analysis. You can configure that by clicking on your user name (on the top right) and doing the following: 1. Clicking on "Generate API Key" and jotting down the generated Webhook URL. PagerDuty will need that. 2. Filling out the "Organization Email" text box. We will send your our analysis there!
Given that we'll be dealing with potentially sensitive data, we reluctantly decided to add a layer of security and have folks register with us first - this allows us to protect your data better. My apologies for the inconvenience.
I'd love to see what the HN community thinks and how we can make it better!
Adding magic into the mix hoping to surface the right thing at the right time but not being entirely sure that would be the case seems like solving a problem of unknowns by hoping that adding new unknowns will cancel out the old ones.
Being an engineer myself, this was a personal pain point and I wanted to solve it, but the key question was whether or not machine learning would help. Thus, most of the time was spent deploying the tech with early adopters, refining the algos, and trying to better understand the value.
What I learned was that our message resonated with some companies more than others. Working with those guys and getting some proof-points on the value is what kept us going!
Business leaders often enjoy seeing their engineers act out this scene. It's gratifying to see someone take expected action, and gratifying to simply look at your direct reports to see if they're working.
But this theater play requires engineers to be idly starting at dashboards first, so they are in the right place to see an issue and take action. This leads to bored engineers, complacency, and delays issue resolution. It's also inefficient to pay people to be bored.
Instead of having a dashboard display a subset of metrics, have alerting configured on these. Pager duty notifies the same if you're in another app, another castle, another room, or another state - you can pay people to do other things instead of staring at a dashboard all day.
A big TV full of metrics is a prop that your actors and engineers will ignore.
It can be intellectually gratifying to use brute-intelligence to "save the day" during an urgent incident, but those heroics are part of the "Star Trek Bridge" scene play, and wastes crucial time.
At a sufficiently large company, looking at dashboards can be difficult because:
1. there are so many dashboards to look at... for so many different services
2. spurious correlations between two graphs showing unrelated events can lead you down the wrong path, if you don't confirm your dashboard-generated hypotheses with log statements or other information.
#2 can be solved by a stronger reliance on logs, http://opentracing.io/ style distributed tracing, and other information.
Something like this Show HN would be useful for problem #1 though.
If the metrics are known as important to your app, monitor specifically on those - that way you're alerted because of low memory or pooled connections or something you're watching, and can lead your team with that info. If the metrics aren't known to be important, why are you wasting your time looking at them on a dashboard?
If you're paged - when it's raining is the wrong time to patch the leaky roof - it's the wrong time to debug and fix a problem in code. Those notes and todos should be pulled into next sprint triage. In the moment, just restoring service should be priority. If restoring service requires modifying code, database, routes, etc, then your testing environments and change control policies need improvement.