Grafana OnCall: an easy-to-use on-call management tool
grafana.com
grafana.com
Stuff like national holiday awareness, integration to vacation calendars, a better UI for swapping days/overrides, etc.
PD schedule checking and trade negotiation becomes yet another thing in the long list of things I need to do when taking a day off. HR system request off, Department Outlook calendar update, PagerDuty coverage check, Outlook out-of-office status & auto-replies, Slack set away, update status AND pause notifications.
I suppose that's because as an on-call developer I am not the user. The user, management who bought the product, gets KPIs & pretty graphs, so they are happy.
(If I were starting a company tomorrow, though, I'd use Linear. Nicest issue tracking tool I have ever seen. It has all the "big business" features like roadmaps and story points, with a lot of friendliness for the individual contributors -- dark mode, keyboard shortcuts, a dedicated triage UI. It's so nice to see someone finally get it right!)
In that it solves problems for one user (project managers) by shoveling it onto another user (developers or admin assistants).
It solves some problems, but it could definitely do a better job of decreasing overall tracking work required, instead of just moving it around.
- Get directions to anywhere on the continent
- Send and receive texts to my friends
- Answer and take a call from a human
But if PagerDuty calls me, Stephen Hawking's speech synthesizer brusquely yells at me and demands I take my hands off the wheel and press a button on my phone to acknowledge the alert. No voice recognition, no ability to kick off an automated play. It's a time portal to 1997! Even the _banks_ have friendlier phone automation these days!
Do you shut down your service for Labor Day? I don't.
I do agree that trading on-call shifts is not very easy within the UI. Part of me dreams of being able to make enough advantaged trades to end up never on-call, like the padre who doubled his holdings in a WW2 POW camp: https://www.ft.com/content/c523efe6-9973-11e1-9a57-00144feab...
The fact that the product seems to have no concept of holidays when its essentially a scheduler++ is a problem.
The problem with doing trades which is the default easiest thing to do given the PagerDuty interface is that when you come back from (or just before you go on) vacation you typically end up with extra bonus on-call shift outside the cycle. Delightful!
All these things just sort of pile up into the "maybe its just easier not to take a couple days off" category, which is not really a mistake on your employers part.
Depends on the service and industry. Banking adjacent companies are often allowed downtime off US business hours. Even at big tech companies I've run internal services that had business hours support only (nonproduction sandboxes, non-business-impacting services, long running job services with SLOs measured in hours or days)
Honestly as a consumer this pisses me off. I get home for the day, relax, eat dinner, and log into the bank to check my finances at 8pm and east coast banks throw up a "scheduled downtime for upgrades" notice.
We've already received multiple questions about OSS and on-premises. Will roll cloud version first, see how it works, collect feedback and build (and share) future plans!
Same time we've focused on making it useful for those who don't use Grafana for monitoring. Feel free to sign up in the Grafana Cloud and use just OnCall if you want.
1. Quickly creating a proper chat 2. Quickly creating an incident document where you can pin chat messages and use it in the post-mortem. Ideally, pinning some graphs that you'd extract from your observability solutions 3. Having a status page to put a small description for non-technical stakeholders.
PagerDuty covers some of this. Monzo's Response [1] and now incident.io [2] try to cover it too. I'd like to have this experience end-to-end.
1 - https://github.com/monzo/response 2 - https://incident.io/
+100 on the creation of incident chat rooms and pinning data to re-use in incident docs. There is nothing worse than copying the timeline events from one tool to a Google Doc.
1 - https://www.indexventures.com/perspectives/incidentio-raises...
> Alerts from the whole team 500 5 minutes
> API requests per API key 300 5 minutes
Product looks great but those API request limits are too low, because alerts rain when you are having incidents and rate limiting all of them is harmful. That's why other products have deduplication keys / aliases so you don't miss important ones.
https://grafana.com/docs/grafana-cloud/oncall/oncall-api-ref...
I'd question the configuration which fires that many alerts in that time frame, and suggest improving alert aggregations and dependencies to get the number down to one or a handful of meaningful alerts.
Also, in my experience with those systems, they only make sense to use very sparingly. Your monitoring becomes extremely fragile when your aggregations and dependencies get complicated enough that "what will our alerting system do when X happens?" results in a flow chart with 18 steps.
If you aren't careful, you can end up making your aggregations less useful than the raw alerts would be.
We just had a short outage where an editor removed the index page in the cms which is central to the site. It's stupid that this is possible but we just operate the cms while we build and operate everything around it for our customer.
I think a large part of our alerts where triggered all at once but the one thing they had in common was that the alerts all pointed to the index page in the cms. E.g. the public www alert for index, the public api alert for index, the preview www alert for index, the preview api alert for index....
Care to link to the docs? I'm interested.
From the article:
With Grafana OnCall’s automatic grouping of alerts within Slack, you can avoid alert storms and reduce the noise your teams are exposed to during an incident.
Seems like the same feature described using different terminology.
What happens if you get 1000 API calls about "Alert 1" and 1 API call about "Alert 2".
You want both on call's to trigger once, but will alert 2 get though?
In reality, probably a lot of missed downtime events, and ops sleeping peacefully I guess.
If that is within your outage model, you'd probably want a redundant on-call service I suppose, even if it's just escalating to a single known email or sms group.
> Once an Incident is triggered, PagerDuty will deliver the First Responder Alert within the Notification Delivery Period for 99.9% of the notifications sent by PagerDuty for the Customer during any calendar month. The “Notification Delivery Period” is five (5) minutes and it is measured as the time it takes PagerDuty to deliver a First Responder Alert to telecommunication providers in accordance with the Service configuration and Contact Information.
> ...
> If PagerDuty fails to meet the SLA set forth herein, Customer may receive a service credit. Customer will be eligible for a credit toward future fees owed to PagerDuty for the PagerDuty Service. The Service Credit is calculated as ten percent (10%) of the fees paid for or attributable to the month when the alleged SLA breach occurred.
If you want to see what teams you are on as the current logged in user, the only way to do it as far as what support told me, is to search for yourself and then check that result.
Disclaimer: Am an employee.
Disclaimer: I work at xMatters.
Their free and lower prices tiers offer a lot of what others have on their top/most expensive tiers. Also, integrations with various alert sources are just easier in most cases. I spent I don't know how long trying to get OpsGenie to work before I gave up.
1. Alerting: Phones you when your servers are down.
2. Incident Management: Help coordinate a response across multiple people.
For the first, there's also: - OpsGenie (owned by Atlassian)
- Squadcast
- VictorOps (now Splunk On Call)
- xMatters
- PagerTree
For the second, there's a bunch of new contenders: - Datadog now has an IM product
- Blameless
- Rootly
- Incident.io
- FireHydrantI've also used OpsGenie (Atlassian now) and really enjoyed it. The amount of integrations they have is staggering.
I may be biased as a co-founder of Spike.sh, but I think we have one of the best designed incident management products out there. We've focused on making it easy to create on-call schedule and overrides, and added templates for escalation, on-call and alert rules.
It's a team calendar to share recurring tasks as a team. Things like PR reviews, who's on support, or who's qualifying leads.
It has far less features than PagerDuty or Grafana OnCall but it serves well a bunch of customers looking for a simple tool to manage team schedules.
We're (more or less) using OpsGenie's free tier, however their scheduling never really "clicked" with me... not sure if i'm special in that regard, however i find the UI/UX pretty... weird...
I need corresponding mobile phone applications for any alert product I intend to use that can override DND/volume etc. on my phone so I can get woken up at night and respond to problems.