Grafana Labs launches free incident management tool in Grafana Cloud
grafana.com
grafana.com
Surprised to hear this as none of our Enterprise plugins had a change in licensing (i.e. going from free to paid, or shifting within paid tiers) as far as I am aware. Would love to dig into this further.
If you're up for it, feel free to send me a note at divy.goel@grafana.com or let me know how best to reach you!
https://grafana.com/blog/2020/07/22/introducing-the-new-and-...
As this is not how we (try to) operate, I also had a look.
From what I could find, it seems the account you are referring to is a very early Cloud account. For reasons I don't know and which might be lost to history, your account had an old and non-standard license attached to. From the viewpoint of today, the license itself seems broken.
To be clear, that is neither your fault nor do I believe you could have caught it even if you looked. While there was no change in the licensing requirements on the plugin, an upgrade to a significant rewrite of the plugin "fixed" the problem of accepting a broken license. That "fix" meant the plugin stopped working.
Again, this is not your fault. But it was not a deliberate action by Grafana Labs nor caught by our testing, either.
I believe your company is in contact with David Dorman, our Head of Self-Service, about this. If you'd like me to ask him to follow up with you directly as well, please let me know how to best contact you.
I'm curious—what other sorts of services are you referring to in your comparison?
Edit: We're happy with incident.io but free is compelling if the product is good and having a single view for observability is useful
happy to pull something together for ya if there are particular workflows you are most interested in comparing
farhan.manjiyani@grafana.com
One thing I'd say is that I find the "react with a robot emoji on slack to add information to the timeline" as a little kluge, hopefully that's not the only mechanism for doing that.
Also does this tool have a postmortem workflow? I didn't see one in the documentation and that seems like an important part of the incident response process.
re: postmortem workflow. The timeline view is built to help postmortems, one of the ways we're doing this is making it easy to paste the info from the timeline as rich text or markdown from the timeline into your post mortem workflow. You'll see that on the top right of the timeline view. We have a lot more ideas in this area and will be investing in this.
Curious if there is specific features you'd like for postmortems?
As your company has shown zero respect for its customers, I will not be using any of your systems.To be absolutely clear it is fair to charge for any of your products, however, if you change it from freemium to paid you can't just pull the plug without reaching out.
The second piece of functionality is around action items. Almost all postmortems generate action items. I need a way to tie action items to a specific post mortem and integrate automatically with my project tracking software (Jira, Asana, Linear, whatever). Ideally there's some top level reporting functionality that shows me the status of post mortem action items.
A nice to have would be automatically scheduling the post mortem meeting with a calendar integration.
On the action item section today actions are tied to Github issues and we are working to extend those - namely Jira/ServiceNow. We do already have a few that gives you a quick look at current status, how long the incident took, and open action items which you can filter by labels or severity (e.g. all severe production issues)
The postmortem doc, meeting link and slack channel are all automatically created for every incident
We've been using Rootly https://rootly.com and love it.
We work with 100s of companies like Canva, Grammarly, OpenSea and others to help build a consistent incident response process on Slack if you're interested. Happy to give you the no-BS sales demo.
FWIW - we are big fans of Grafana and have a native integration (think automatic Grafana metric/dashboard snapshots into #incident channel.
Bonus questions, are you tracking or driving improvement in the related times for detection/response/mitigate/recover?
Disclosure: Principal at AWS currently in a similar apace. Though I ask in a personal capacity and interest.
I'd be very interested to hear your thoughts too?
In short yeah, using the incident status as the implied times makes sense for the bulk of cases. Totally agree on picking out signal from the users inherent actions, but allowing them to provide more specific data when they know better.
Digging in a little further Im personally interested in moving past the incident data and inspecting the incoming alert(s) and related telemetry/metric/alarm data. For example think of the alarm definitions like “five 1m datapoints with a value above 0.1.” There’s a good argument to count impact (and incident duration) from that first datapoint > 0.1. Then theres the delta from metric processing to alert to incident creation. On the backend theres frequently a delta between mitigating impact and actual incident resolution, again I think getting back to the source alarm/alert/metric data would get us a more accurate view of operations and customer impact.
Incident is available in the free tier, that's awesome... are there any limitations on that at all? Is the free tier version of Incident as fully featured as the paid tiers?
I wrote a little glue to make this straightforward for anyone else who uses Prometheus/Alertmanager: https://github.com/jrockway/alertmanager-status This ensures that the website check checks the health of the whole alerting pipeline; Prometheus has an always firing alert, Alertmanager is set to send that alert to alertmanager-status, and alertmanager-status starts failing its external health check if it isn't seeing that alert firing at the configured interval. If one of [Prometheus, Alertmanager, alertmanager-status] fails, then your website health check fails.
FWIW our core application is hosted in AWS but we maintain our own Grafana infrastructure independently. So it's not hosted on our own tech, per se, though we're still responsible for keeping it online.
This looks great & would also love to see a self-hosted option. Honestly the more stuff like this that gets rolled into the OSS Grafana it actually makes me both more likely to try it and then more likely to eventually end up on the managed Grafana Cloud, as I will inevitably get sick of trying to maintain our own separate infra & the business case for centralising in Cloud makes more and more sense.
We have plans to build a native mobile app for ios & android for OnCall that would let you achieve this over the next few months.
OnCall is a separate product from Incident. It's available via OSS and Cloud. Incident and OnCall work well together, or you can use either as standalone!
Internally we (Grafana Labs) already have replaced PagerDuty and are using it for our teams running critical systems.