Anyway, this is at least the 7th or 8th time Slack has been totally or partially down during working hours.
This is very, very far from "rare outages".
We can have the best of both worlds.
I just don't see it being worth the extra effort of self-hosting a redundant copy of the system, and then keeping that system up to date with the latest software for security patches / features / bugfixes. Especially when the cloud version has n x 9's of uptime.
>If we fall short of our 99.99% uptime guarantee, we’ll refund customers on the Plus plan 100 times the amount your workspace paid during the period Slack was down.
Source: https://get.slack.help/hc/en-us/articles/204113126-Plus-plan...
Maybe an emergency mode in the Slack application. When users cannot reach the Slack servers (when they're down or bad network connectivity) the application exposes a P2P connection and looks for other P2P instances.
It would still need a centralized tracker (like a torrent tracker) to help P2P connections find each other through the internet.
All this is easier said, then done!
Why have a mothership at all? Why couldn't software clients simply talk to each other? We already have hash-based addressing systems that can find known peers anywhere in the world (e.g. Tor).
I'm also pretty sure all of us wrote this in CS 101. I sure did.
We manage over 900 virtual systems on many, many hosts and file systems and have next to no downtime because we work really hard to make that the case.