I think we should push for a metric where "up" means 100% of people that want to use the service are able to use the service. If 1% of users can't send messages, then that should count as a full-blown outage and should start counting against whatever SLA they advertise.
The underlying problem here is that apparently everyone lies about uptime, so if you don't, that looks bad to potential customers. I fear that we will have to push for some legal regulation if we want accurate data, and ... people will probably be opposed to that.
I was wondering why the link from a Jira wasn't opening in slack, the page eventually timed out and gave me a link to status.slack.com where it told me everything was peachy. Cue me wasting time trying it again because apparently there was no issue with slack..
Honestly, I could live with a 99.50% SLA, if that's what it really was. After today's probably full-day outage, they'd just have to be extra careful for the rest of the year (or pay me money). Kind of sucks when it's 1/4 that you blow your year's SLA budget though.
I mean, that’s nice to say, but how do you measure/prove it?
Certainly, having the SLAed party check themselves is silly. But what are the other options? If it was up to the customer, customers could make up faults to get free service. (Since it’d be up to the customer to prove, and customers are generally less technical than vendors, you’d have to expect/accept very non-technical — and thus non-evidentiary! — forms of “proof”, e.g. “I dunno, we weren’t able to reach it today.” Things that could have just as well been their own ISP, or even operator error on their side.)
IMHO, contractual SLAs should be based on the checks of some agreed-upon neutral-third-party auditor (e.g. any of the many status/uptime monitoring services.) If the third party says the service is up, it’s up in SLA terms; if the third party says the service is down, it’s down in SLA terms.
(And, of course, if the third party themselves go down, or experience connectivity issues that cause them to see false correlated failures among many services, that should be explicitly written into the SLA as a condition where the customer isn’t going to get a remedial award against the SLA, even if the SLAed service does go down during that time. If the Internet backbone falls over, that’s the equivalent of what insurance providers call an “act of God.”)
But in a neutral-third-party observer setup, you aren’t going to get 100% coverage for customer-seen problems. An uptime service isn’t going to see the service the way every single customer does. Only the way one particular customer would. So it’s not going to notice these spurious some-customers-see-it-some-don’t faults.
So, again: what kind of input would feed this hypothetical “100% of customers are being served successfully” metric?
ETA: maybe you could get closer to this ideal by ensuring that the monitoring service 1. is effectively running a full integration test suite, not just hitting trivial APIs; and 2. if gradual-rollout experiments ala “hash the user’s ID to land them in an experiment hash-ring position, and assign feature flags to sections of the hash ring” are in use by the SLAed service, then the monitoring service should be given N different “probe users” that together cover the complete hash-ring of possible generated-feature-flag combinations. Or given special keys that get randomly assigned a different combination of feature-flags every time they’re used.
Google published a paper last year describing this approach to measuring uptime: https://blog.acolyer.org/2020/02/26/meaningful-availability/
The idea is to define availability as "the probability that the site 'appeared' to be down for a random user, averaged over a time window of size w". You can choose a particular value of w and look at trends over time, or you can plot availability as a function of w to understand patterns of downtime.
If you look at the history page you can see its not 100% for every month: https://status.slack.com/calendar
But it was roughly "one large impact a month, for six months", with large caveats that upper management for whatever company had to be working with the product during that month.
Large companies don't care if X service went out during the night and impacted someone not in their timezone.
If the CTO notices that he can't use something with the same regularity that he gets paid, then it doesn't take long for it to stick in their mind. But migrating everything is _so painful_ that the majority of large companies will do anything they can to avoid moving away.
This is a key point is the popularity amongst VCs in investing in B2B SaaS. I take their (and your) word for it. But honestly, I don't actually understand this.
Why is migration so hard?
This is not to mention the fact that half our staff aren't hugely technical, so have actively _learnt_ how to use Slack and it's features around notification control (things that may come "naturally" to the tech-savvy crowd on HN), @-things, bots, etc, and they would need to re-learn a new tool that is going to work in a different way.
This would be a substantial effort for us, and we're a small company. Are there ways to materially minimise this cost?
Why would anyone make it easy
Things that are easy to migrate from get replaced by things that are hard to migrate from, eventually.
IRC is incredibly easy to migrate from.
The parent meant a law as in "a law of physics", not a piece of legislation.
Plus, it will just take a long time to get everyone on board and using the replacement system. My department is slowly plodding towards using Teams over Slack, but there are enough hold-outs (my sub-department being one of them) that it still doesn't have wide-spread adoption.
If you had a small shop with a dozen tech-savvy people and Slack became a problem which was used exclusively for quick business chats, you could probably push a change to another chat platform the next day. You might struggle when you have thousands of employees, some that needed training to use Slack and still aren't that proficient.
* Getting out of the Enterprise Contract, or waiting for the year to end. * Training people on new software. * Loss of productivity. (1) Learning a new UI, processes, workflows -- both individually and organizationally. A feature or concept in "Tool A" may exist in a completely different form in "Tool B". Or not exist, and then people need to adapt to and work around the missing feature. (2) Missing out on needed information due to the above. Ultimately, software exists to move and transform data, and when you change the software people have to adjust. Sometimes that doesn't go great. "Oh, I didn't realize I needed to check this checkbox".
Another way to say this is "organizational inertia", which is a fancy term that means "it's hard for people to adjust to change".
And you might think developers and other technical people would have an easier time of it. They (we) do, but not to the extent you may expect. I've been on the front lines of a handful of migrations that affected only the IT staff, and it was a long and arduous process each time.
Man it bothers me so much when applications change their UIs on updates for no apparent reason other than "it looks better".
IntelliJ changed the way build and debug buttons looked in some update and it took me days to get used to it and I could find them in a snap again. Slack did a couple of no-reason changes as well.
The really big one, for companies of a certain size / cash flow, is compliance. Companies spend a lot of time developing compliant work flows around a service like Slack.
Migrating to another service requires rewriting the compliance narrative. The current compliance people might not have the confidence or willpower to do that effectively, and can raise legal objections to any such migration indefinitely.
* work on code
* update JIRA
* complete required trainings
* work on my peer reviews (Workday/Okta are up)
* review tech specs
Deploying is actually a very small part of my job.There's also this perverse incentive to Slack all the things. Lots of CI notifications are sent through it. Some org processes are implemented as workflows. There's been talk of how wonderful it would be to hook up tasking and work tracking to slash commands. I and others often use Slack instead of the 'official' tool to video call each other.
An outage like this is still really disruptive. It's not like everyone realizes what's going on immediately or at the same time; we have backup tools, but our turn radius is pretty wide. Some of us can't even communicate effectively without memes, too, and backup tools don't have a giphy integration.
EDIT: Do your CI integrations fail if Slack can't be contacted? Do those failures fail your pipeline? Whoops!
However, I drive everything through slack - GitHub, linear, calendars, Notion, support emails, etc. I have notifications turned off for every service we use except for slack. This allows me to effectively ignore everything except for slack. These types of outages destroy that workflow for me.