Azure, Teams, Outlook are almost down from Greece and Germany, and their status page shows that everything is fine :-)
Azure, Teams, Outlook are almost down from Greece and Germany, and their status page shows that everything is fine :-)
It's about contractual obligations and SLAs. Things are not officially down in most agreements until MSFT acknowledges they're down. Refunds issued because your blob storage failed to meet 99.9999 uptime to your largest customers are directly tied to these statuses.
I think it's an important enough page that it can't be automated. It needs a manual approval from a human, for the very basics, like even if the status reporting system is operating correctly, because of various downstream effects.
If it's a false positive they just resolve it without it affecting SLA and if it's a real problem then us customers wouldn't have to debug our own stack for 2 hours before Microsoft informs us that they are the problem.
EDIT: Wonder how many man-years of extra debugging work their non-working status page have caused the customers.
And so updates to the status page become political and locked behind senior management approvals.. like AWS.
Works equally well. See the point?
(1) The monitoring system would be altered to ignore tests that return false positives (at the expense of missing the alert when it represents an outage).
(2) Fixing the monitoring. It wasn't working for the sysadmins/operators, anyway, since it had so many false positives that their "mental model" was essentially based on (1), anyway.
At least, where I've forced the issue of doing just this, that's exactly what happened. At the end of the day, especially since SLAs took a hit and that affected bonus payouts, monitoring got a lot better -- as did overall team function when we truly realized how bad things were -- we stopped doing workarounds and started fixing problems at a more fundamental level which led to SLAs that were both accurate and excellent.
It helped bring attention to a hidden problem which resulted in time being allocated to fix tests that dropped constant false-positives and to evaluate each for whether or not it should exist in the first place.
It's weird how slow they are with manual sign-off though.
If you work with Microsoft, you might as well spend a few bucks extra and have an external monitoring system monitor Microsoft's systems so you get real-time third-party confirmation when your monitoring alerts you of issues concerning your system. It's the price you pay for scale, I guess. More money involved = more lawyers involved = more accountants involved = more MBAs involved = more corporate bullshit.
I disagree. What if you're having issues and the status page is incorrectly reporting an incident? It would be easy to waste a load of time waiting for the status page to sort itself out, only to find out you've still got an issue.
Thats not the goal.
> It's hard to see how the goal here could be anything other than trying to add plausible deniability for what would otherwise be obvious deception
Thats the goal. The "status page" is considered the source of truth for most of the big contracts. If status-page=OK then your contract with them isn't violated. So changing the status page is a big deal, with real financial implications. The status page isn't a view into the SRE's tickets, its a declaration that the service isn't being provided.
What might be happening is that there is fine print you have to read and be in compliance with in order to be eligible for the SLA.
For example, look at all the conditions which have to be met for a breach of VM SLA in Azure:
https://azure.microsoft.com/en-us/support/legal/sla/virtual-...
Hidden in the SLA details is typically hints on how you can become more resilient in the cloud. So it pays to read the SLA details and really deeply understand what they are telling you.
I'm joking, but...