You won't have access to the things slack integrates with but at least you can talk.
I know some people use IRC as a low-dependency fallback channel.
I had an ircd of some description installed, configured and ready for use in under 3 minutes, and started to get my team using it as a stop gap measure. Had that rolling long before either the outage was resolved, or the HipChat server was resurrected. People forget just how stupidly easy it is to set up IRC.
Freenode runs this one for a server.
[1] - https://www.unrealircd.org/
[2] - https://www.anope.org/
[3] - https://thelounge.chat/
However, when major backbones go down it takes it a while to recover.
[1]: https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...
We build services on top of that that don't cope with failures.
At the end of the day, that sounds like a massive pile of SPOFs that just barely keep up..
As app is 7 but we have grown beyond the seven into meta layers of service types which have become the favric of how we operate... such as email, txt, chat etc to keep all other aspects of infra up and running.
Also, everything above layer 4 is a enormous mess, considering that stuff is usually happening at the application itself anyways. There are network protocols which have proper OSI seperation, but they are rarely used. (IS-IS in combination with CLNS for instance).
A failure of mine was to not keep a project portfolio runbook which i could look back upon.
Imagine a system where all your status reports were logged and searchable - that would be move valuable than linkedin...
Sure, having a single connection is a SPOF, but anyone living in an urban area and anyone with a cellphone (and not using the same cellular carrier for home internet) has more than one connection.
A lot has to go wrong for me to not get work done. If Slack goes down, I use text, email, or Jira comments (in roughly that order of priority). Relying on only one communications method is foolhardy.
I had a similar outage at another job for a day. Pre slack days, but instant messaging and email was down.
I needed to do some things so I just stopped checking with people. So did other people.
The result was I found was that we were double checking, coordinating and doing a lot of verification that honestly wasn't needed. The sky didn't fall, everyone still saw the changes and were ok with it.
After that folks stopped doing a lot of the traditional coordination that proved to be superfluous and maybe never accomplished anything. The handful of mistakes that happened, were also easily caught / fixed.
Granted, this requires people to make good choices.
If Slack is your only forum of communication with your team it’s time to rethink your support structure and DR plans.
Edit: It’s very unhelpful if I point out issues and don’t provide solutions. PagerDuty has a great article about Incident Response[1].
Ideal would be a live conference call (phone, Skype, Hangouts etc...), but you've got SMS, email, heck, I once solved a pressing issue with my team in a Sharepoint Word document where all our comments were just new paragraphs being edited live (which actually had the handy benefit of being savable without needing to pay extra for the privilege...).
* do not fit well for a quick message interchange like you can have during an outage (email)
* don't have the whole company (or at least the whole tech team) on it (whatsapp, telegram, you name it)
* make compliated to share big chunks of text/logs (phone call)
It probably makes sense creating another secondary - usually dormant - communication hub like a WhatsApp group for emergencies but people tend to misuse it (like using it to ping for some minor incident when all the other more suited communication tools are working perfectly).
Anyway it wasn't a big outage on our systems, just a minor hiccup, but it's a lesson learnt nonetheless.
If there was a lower reliance on IM for all tasks, other solutions for crisis management would be utilised (potentially).
> failing to address pending scalability issues
How is that not related to not getting enough funding and resources?
And yes, having scalability issues means moving too fast - the _business_ moving too fast.