Slack uses its own product for all coordination during incidents and outages, so what happens when it goes down? “To the best of our ability, we use Slack,” says Richard Crowley, Director of Operations at Slack, “[but] If we’re really having a bad day, we use Skype.”
I was annoyed when after a several hour long outage SendGrid still had not updated their status page to reflect an issue.
But I was even more annoyed when 5+ hours into the outage, a different support person told me they aren't aware of any outage because their status page does not reflect one.
If I sound bitter it's because I am, I was the on call person who got woken up in the wee hours of the morning to deal with this. I'm referring to the outage you guys had in December if you're wondering.
They were rather rude and kept insisting that it wasn’t an issue on their end. Eventually we gave up and migrated away from SendGrid. If I remember correctly, they did fix it, at least partially. I think it had something to do with one of their prod servers running a dev config, resulting in DKIM signatures for a SendGrid-owned test domain—something along those lines. That may have been a different incident, though.
That was just the straw that broke the camel’s back. There were often issues that only affected one or two servers—we were often told that a server didn’t have an up-to-date configuration or something like that. Each time, it was the same process: convince them it wasn’t an issue on our end, escalate it repeatedly until it lands on an engineer who understands the issue, they fix it in an hour after wasting days of my time insisting that yes, I know what a DNS record is. And the whole time, some (small) percentage of emails are failing. I can only provide so many logs, DNS records, and email headers to demonstrate the issue; at some point, someone has to be able to look at it and go, “well, yes, that’s very obviously the wrong value, and nothing a customer sets on their end should ever have that effect.”
Sure, Slack dropped the ball here, but at least they updated their status page. Even for small issues, when I contact them, they either get it resolved or tell me it’s on their to-do list—no need to convince people that I understand basic concepts.
At the very least they should be hosted isolated of your production service components at all layers, so both don't go down at the same time.
For slack my guess is their internal slack is not hosted on slack prod and hopefully also does not share components below L4 as well ( DNS, IP , domain, routing etc )
Fair, I guess.