If Slack is your only forum of communication with your team it’s time to rethink your support structure and DR plans.
Edit: It’s very unhelpful if I point out issues and don’t provide solutions. PagerDuty has a great article about Incident Response[1].
Ideal would be a live conference call (phone, Skype, Hangouts etc...), but you've got SMS, email, heck, I once solved a pressing issue with my team in a Sharepoint Word document where all our comments were just new paragraphs being edited live (which actually had the handy benefit of being savable without needing to pay extra for the privilege...).
However, when major backbones go down it takes it a while to recover.
[1]: https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...
We build services on top of that that don't cope with failures.
At the end of the day, that sounds like a massive pile of SPOFs that just barely keep up..
As app is 7 but we have grown beyond the seven into meta layers of service types which have become the favric of how we operate... such as email, txt, chat etc to keep all other aspects of infra up and running.
Also, everything above layer 4 is a enormous mess, considering that stuff is usually happening at the application itself anyways. There are network protocols which have proper OSI seperation, but they are rarely used. (IS-IS in combination with CLNS for instance).
A failure of mine was to not keep a project portfolio runbook which i could look back upon.
Imagine a system where all your status reports were logged and searchable - that would be move valuable than linkedin...
Sure, having a single connection is a SPOF, but anyone living in an urban area and anyone with a cellphone (and not using the same cellular carrier for home internet) has more than one connection.
A lot has to go wrong for me to not get work done. If Slack goes down, I use text, email, or Jira comments (in roughly that order of priority). Relying on only one communications method is foolhardy.
If there was a lower reliance on IM for all tasks, other solutions for crisis management would be utilised (potentially).
> failing to address pending scalability issues
How is that not related to not getting enough funding and resources?
And yes, having scalability issues means moving too fast - the _business_ moving too fast.
* do not fit well for a quick message interchange like you can have during an outage (email)
* don't have the whole company (or at least the whole tech team) on it (whatsapp, telegram, you name it)
* make compliated to share big chunks of text/logs (phone call)
It probably makes sense creating another secondary - usually dormant - communication hub like a WhatsApp group for emergencies but people tend to misuse it (like using it to ping for some minor incident when all the other more suited communication tools are working perfectly).
Anyway it wasn't a big outage on our systems, just a minor hiccup, but it's a lesson learnt nonetheless.
You won't have access to the things slack integrates with but at least you can talk.
I know some people use IRC as a low-dependency fallback channel.
I had an ircd of some description installed, configured and ready for use in under 3 minutes, and started to get my team using it as a stop gap measure. Had that rolling long before either the outage was resolved, or the HipChat server was resurrected. People forget just how stupidly easy it is to set up IRC.
Freenode runs this one for a server.
[1] - https://www.unrealircd.org/
[2] - https://www.anope.org/
[3] - https://thelounge.chat/
I had a similar outage at another job for a day. Pre slack days, but instant messaging and email was down.
I needed to do some things so I just stopped checking with people. So did other people.
The result was I found was that we were double checking, coordinating and doing a lot of verification that honestly wasn't needed. The sky didn't fall, everyone still saw the changes and were ok with it.
After that folks stopped doing a lot of the traditional coordination that proved to be superfluous and maybe never accomplished anything. The handful of mistakes that happened, were also easily caught / fixed.
Granted, this requires people to make good choices.
More study are needed to clarify if this work-related communication might actually be required for work to get done. (Cal Newport is definitely preparing a book on that any time soon.)
Unfortunately, the researchers are too busy maintaining their IRC server to actual do it - but at least they're using IRC.
Whereas, my productivity definitely goes up when HN is down ;)
* update the version of the IRC server
* update the version of the os of the machine running the IRC server
* repair a broken disk / fan / overflowed disk on the machine
* add / remove / reset password / change weird settings of users
I'm not saying this is "unbearable", and plenty of organisations have people whose job description would probably correspond to doing those tasks.
But I can definitely understand why you would want to skip them entirely and have someone host your chat - which is basically the job slack is paid for.
"Mission critical, but not paid by your customer" is always tricky to staff for, isn't it ?
Odds are that you’ll have updates worth installing once every couple of years.
>* update the version of the os of the machine running the IRC server
Almost never unless there’s a remotely exploitable code execution vulnerability. Local bugs wont matter unless you run multiple services on the same box.
>* repair a broken disk / fan / overflowed disk on the machine
Depends on your hosting setup. With a cloud setup perhaps never.
I’m certainly not trying to suggest that anyone should use IRC over slack in any situation, just that it’s not a horrible time sink requiring significant maintenance.
In this forum you'll be hard pressed to find someone who thinks it's a good idea to leave a cloud-based (or any internet-facing, really) server completely without maintenance for any significant period of time.
If you had said that usually it's as simple as checking in every now and again to reboot / make sure security updates are installed / check functionality and that it's a manageable burden which is worth the effort you might have gotten a more favourable response.
Edit: or if you had automated it to an acceptable level, maybe some more information about that.