Slack was down
status.slack.com
status.slack.com
Outages mean down, badly, and require RFOs. In my experience working at an ISP we only did any sort of "blame-game" if customers demanded an RFO. We still did post-mortems but usually in a more gentle, roundtable way.
Setting a goal of "Zero or Near" down times doesn't necessarily mean less down times, but it usually does mean politics, shell games and 'clever' bookkeeping by middle and lower management, all of which is clearly visible to the rank and file.
Before you had down times and customers wondering if they could trust you. After a goal of "Zero or Near", you still have "down times", and you still have customers wondering if they can trust you while they stare at a status page is clearly lying to them.
But now you have employees watching their managers lying to their bosses, and wondering why they shouldn't do the same. You have a fall in morale, because those pretty numbers usually come at the expense of the things on the "Our Values" page. But to the Suits the graphs look better, and that's the important bit.
This isn't how is always is or always goes, but I've seen it enough.
Your goal is to identify service impacts, do post mortems where no one is afraid to say something went wrong, and to improve processes to prevent future events. You can't do that if there is a culture of blame.
My comment was more of a dig at Github specifically. Recently things haven't been working for me (from my perspective an outage) and when you go to their status page it just says "Degraded performance". When your performance has degraded to the point of failure it's time to acknowledge that is in fact an outage...
Messaging and Connections.
- Login/SSO
No issues
- Connections
Something's not quite right View details
- Messaging
No issues
- Link Previews
No issues
- Posts/Files
No issues
- Notifications
No issues
- Calls
No issues
- Search
No issues
- Apps/Integrations/APIs
No issues
- Workspace/Org Administration
No issues
Perhaps I should have been less subtle than opening with a 'Heh', and made it clear the claim that only the messaging & connection parts were failing was tongue-in-cheek.
This page's title is "Slack is down" but Slack's status page at the time was modestly conceding that "Something's not quite right".
Page asserted an 'incident' only, not an outage, but for all intents and purposes a failure of the 'messaging' component for a messaging tool is effectively a full, major, total outage as far as your users are concerned.
At the very least they should be hosted isolated of your production service components at all layers, so both don't go down at the same time.
For slack my guess is their internal slack is not hosted on slack prod and hopefully also does not share components below L4 as well ( DNS, IP , domain, routing etc )
Slack uses its own product for all coordination during incidents and outages, so what happens when it goes down? “To the best of our ability, we use Slack,” says Richard Crowley, Director of Operations at Slack, “[but] If we’re really having a bad day, we use Skype.”
Fair, I guess.
I was annoyed when after a several hour long outage SendGrid still had not updated their status page to reflect an issue.
But I was even more annoyed when 5+ hours into the outage, a different support person told me they aren't aware of any outage because their status page does not reflect one.
If I sound bitter it's because I am, I was the on call person who got woken up in the wee hours of the morning to deal with this. I'm referring to the outage you guys had in December if you're wondering.
They were rather rude and kept insisting that it wasn’t an issue on their end. Eventually we gave up and migrated away from SendGrid. If I remember correctly, they did fix it, at least partially. I think it had something to do with one of their prod servers running a dev config, resulting in DKIM signatures for a SendGrid-owned test domain—something along those lines. That may have been a different incident, though.
That was just the straw that broke the camel’s back. There were often issues that only affected one or two servers—we were often told that a server didn’t have an up-to-date configuration or something like that. Each time, it was the same process: convince them it wasn’t an issue on our end, escalate it repeatedly until it lands on an engineer who understands the issue, they fix it in an hour after wasting days of my time insisting that yes, I know what a DNS record is. And the whole time, some (small) percentage of emails are failing. I can only provide so many logs, DNS records, and email headers to demonstrate the issue; at some point, someone has to be able to look at it and go, “well, yes, that’s very obviously the wrong value, and nothing a customer sets on their end should ever have that effect.”
Sure, Slack dropped the ball here, but at least they updated their status page. Even for small issues, when I contact them, they either get it resolved or tell me it’s on their to-do list—no need to convince people that I understand basic concepts.
It's crazy how easy it is to depend on Slack. Email collaboration feels downright archaic. What a great product!
What is everyone using?
Getaether.net/pro
E-Mail is good enough to let people know that something like Slack is down. It's close enough to real-time to be used as a backup for something like this, and everybody already has an e-mail address.
My company already has Slack, HipChat and Teams, and it's a communications nightmare. People are constantly confused and asking for which chat client one manager uses, while others don't know where to go for the support channels.
Maybe this sounds like old-school IT, but the new-school "set everything up ad-hoc using your personal shit" just sets you up for disaster down the road.
That being said, I don’t think I’d want Telegram to be my failover. They have a tendency to experience regional issues with great frequency, and if you’re in a lot of active chats, things start to get buggy.
I'm curious about how many people in the company are affected by these problems. Most developers and infrastructure people use the same communication tools, but perhaps support and sales may use their own tools. Given that, how often do developers and infrastructure need to communicate in real time with sales or support?
I would suspect that it's not often enough that email wouldn't work as a way to handle issues or open support tickets. Now managers have to deal with communications between their own team and other teams and parts of the org, but, in my opinion, having to deal with different chat platforms is part of their job and they should be the ones to deal with it rather than forcinng everyone to use the same platform, even if it's suboptimal for certain users.
See https://status.slack.com/2020-05/87ad1d12e36fcf0c for incident details:
> [May 12, 4:53 PM PDT] Users have reported general performance issues such message sending failures and timeouts. We’re working to get things back to normal as quickly as possible and will provide an update shortly.
Also the lag and human intervention is by design as there will be a lot of false positives as some error rates in these systems are expected, a very small percent of API calls will result in errors always like unhealthy instances etc. Unless the issue is substantially large it not in user(alert fatigue) or the business interest to declare it down.
https://en.wikipedia.org/wiki/Glitch_(video_game)
https://techcrunch.com/2008/04/02/game-neverending-rises-fro...
"We are currently experiencing an outage on our network. We are working to resolve this as quickly as possible. Please check http://status.cogentco.com for updates on this event."
503 Service Unavailable No server is available to handle this request.
And thanks to the electron-ness of the app, on the mac you cannot even drag the window - the titlebar is drawn by javascript and without it, it cannot be dragged