SLO alerting for mortals
ervinbarta.com
ervinbarta.com
This article assumes relatively constant usage. If a service is used heavily in a single time zone this number could be considerably lower.
The flip side of that is that it might not make sense to alert during night hours of that time zone. You wouldn't want to get paged by a single user spaming retry on an error page.
I currently want to do that for a service built on several cloud services with their respective SLAs and my approach is to go through the combined probabilities to get an effective error rate. That‘s a bottom-up approach. I‘d combine that we a top-down derivation of what SLOs are required from the business side. If the first number doesn’t fulfill the business requirements with some buffer, we‘ll need to redesign. How do others do it?
—-
SLA: service level agreement, values of KPIs promised to customers
SLO: service level objective, internal target values for those KPIs, typically slightly more demanding than the SLA
SLI: service level indicator, measured values of the KPIs to check against SLOs/SLAs
But in the meantime you can get slick ci/cd, unit, integration, end to end and performance tests. Maybe feature flags, containers and orchestration, with failover, self-healing, global distribution, DORA metrics again, monitoring, alerting, dashboards, visualisations - and then you’re ready for hello world. Or find features customers like.
In such a case, I would encourage you to set the SLAs relatively lower, at least until you can gain some real knowledge from actual customers.
You might also want to approach some prospective customers and bring them into a private preview version (protected by NDAs), so that you can start gathering that data earlier.
Also, ship early. Ship an MVP earlier than you thought possible. It’s okay if it’s all held together with spit and bailing wire, at least you’ll be able to start gathering data sooner, so that you can start the real business of actually building what the customers truly want.
If anyone else has come up with better metrics I'd love to hear about your feedback.
Step 2: Evaluate how much money you would earn by increasing uptime, and evaluate how much the processes required to do that would cost. Incrementally raise SLO levels as long as you benefit from it.