Let's use real words that tell the listener or reader what we're talking about.
Let's use real words that tell the listener or reader what we're talking about.
More correctly, it's something of an unattainable goal. Can we get rid of on-call? Probably not. Can we drastically reduce the frequency and scope of on-call actions? Absolutely! Can we get to where every single time the on-call person has to take action, there's a post-mortem to understand what went wrong, and fix the automation and monitoring so it doesn't happen again?
Have you ever gotten paged for a transient error that went away by the time you checked it? Or paged for the same error 5 weeks in a row? These, and really most pages, are fixed by DevOps.
In a DevOps system, a feedback loop is used to address all alerts as bugs to be immediately and permanently fixed. After a while, if something breaks, it's because someone just changed something, so it's happening during working hours. Alerts pop up in slack and are acknowledged before a page is sent out. So nobody is getting called.
If your infrastructure is ephemeral and managed as code, you use CI/CD to deploy all changes, and you aren't resource-constrained (at this point, cloud-native infra is only constrained by budget) you shouldn't have stuff crashing randomly at 3am, so there should be very few pages.
Isn't "budget" just another type of resource?
Of course Ops is still needed but what they do has probably changed a lot in this case. And their responsibilities and how they work with Dev must changed.
1) let's fire all the DBAs
2) let's fire half the operations team, and outsource the rest
3) let's have the devs do the operational work
What could go wrong?!
Anyone have the German translation?
That's the Americanized acronym of the original German word.