Slack is down
status.slack.com
status.slack.com
https://news.ycombinator.com/item?id=25632346&p=2
https://news.ycombinator.com/item?id=25632346&p=3
(Yes, these comments are an annoying workaround. Their hidden agenda is to goad me into finishing some performance improvements we're badly in need of.)
My bet is that this incident is caused by a big release after a post-holiday "code freeze".
Doubt anyone releasing big changes Monday morning.
Guess it was Slack being Slack.
> Doubt anyone releasing big changes Monday morning.
This is definitely an engineering best practice, and by best practice, I mean something that Uber's, I mean Slack's SRE team strongly pushed for, and got politely overruled on. After a code freeze is lifted, it's quite common for lots of promotion-eager engineers to release big changes.
EDIT: Guys it was a joke, chill
Releasing 1 change a year with a 100% chance of working -- no promotion for 10 years
Releasing 10 changes a year each with a 10% chance of breaking something -- 1 in 3 chance of promotion in a year, and a 2 in 3 chance of downtime
Where did that assumption come from? Also are you claiming that it takes 10x more time to release a non-breaking change?
I don't see any need to deploy a big change at once in the software world today. At worst feature gate the thing you want to do and run it in a beta environment, but still push the actual code down the pipeline.
Every Uber/ex-Uber engineer is nervously chuckling at this comment right now
Did you mean that literally? E.g. is it common at Uber that engineers can release changes to production on their own?
The engineering team is responsible for the mess caused by a bad deploy, so it's appropriate that those engineers should also choose the timing.
Our team typically deploys between 10am and 4ish, local time, since that's when we're at our desks and ready to click through the approvals and monitor the changes as they go through our pipelines.
The feature enablement happens through an EFT / beta process, and the final timing of GA enablement is a PM decision. But features are widely used by customers ahead of that time, as part of the rollout process.
Our team usually rolls out non-feature changes to services via dynamic configuration switches, so that we can get new bits in place, and then enable new behavior without a redeploy. This also enables us to roll back the dynamic config quickly if something unexpected happens.
(We generally don't do this for net new functionality; there's lower risk in adding a new REST endpoint etc. than in changing an existing query's behavior or implementation.)
Although, we also don’t close the pipeline for just any holiday break. In fact low holiday traffic is a good time to keep pipelines open, since changes will impact less people.
That sounds pretty early to think somebody on the west coast did something, other than maybe acknowledge the pages and declare the incident.
- high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, the higher the divergence world of production and the world of the new feature
- lots of new hires (new year = new hiring budget). New hires are missing some tribal knowledge about the system and make a production-breaking release.
I tried to think of other reasons, but these two overwhelmingly stand out as the two biggest reasons. Would love to hear from others.
Hate on the cloud all you want, but AWS has (several flavors of) load balancers and various ways to automatically scale up and down resources (and if you're conservative, you can disable the 'down' part). If you're operating a major SaaS company like Slack and not taking advantage of them, something's gone wrong.
In 2021, how does one keep track of resource starvation at the process, container, os, service, pod, cluster, availability zone and region levels?
January is busy for recruiting, but given a week or two of interviewing and negotiating, two weeks notice, it's probably February before new employees are starting, and they're not making big, production-damaging deploys for a week or two after that.
Probably not as big of a rush as the end of school year rush in summer though.
I also doubt that new people will be breaking production on day one. Even at a fast moving startup I'd expect it to take a bit to go through the onboarding paperwork, get a laptop and actually try pushing a change to production.
People came back to work, and most of them start around the same time (US wise at least).
Hence kids - a vital lesson for all of us - don't start the call at a full hour, give it 3-7 min to make your coworkers confused and give some time for the systems to auto-scale ;)
Likewise, incognito mode will also ignore most cached web content, meaning all assets on the Slack web app will get loaded again from scratch. This "clean state" start could, theoretically, get around issues with old - potentially incorrect/outdated - assets being loaded, even though that really shouldn't happen under most circumstances.
What it could be is some engineer somewhere coming in after the holiday, noticing a slightly flaky thing, and thinking, "I'll reboot/redeploy/refresh this thing so the flakiness doesn't get worse". Only it turns out the flaky thing was a signal of something else falling over. Or maybe the redeploy was the wrong version because of bad CI/CD, or maybe the person just fat-fingered it.
At least that how it worked at one FAANG
(I'm not making a statement if that's good or bad or if it works or whatever. Please don't read an opinion into it.)
Now, caveats etc, this was a collection of single applications in a big microservices architecture, and as the project grows it becomes more and more difficult to manage something like this, especially if you get more pull requests in the time it takes to do a build. But it is the way to go, I think.
Anyway, since tests and CI are not definitive, you also need a gradual rollout - 1%, 5%, etc - AND you need a similar process for any infrastructure change, which gets more and more tricky as you go down to the hardware level.
I'd like to, but unfortunately I don't feel like I can in good faith. Matrix is woefully immature, and suffers from a lot of issues, but I think is closer to being a functional Slack/Discord alternative. XMPP is much more mature, and works very well for chat, but doesn't have a nice package that does all the Slack stuff--at least not that I'm aware of. I'd love to be proven wrong there. I know it can be done, but if it can't be deployed quickly by an already overstressed team member, what chance does it have?
Mattermost is certainly not what I meant. That's just trading one Slack for another.
From my experience, when I self-host stuff it's a lot faster (more server resources) and never had any downtime (server doesn't simply go down for no reason).
We ended making the switch and committed to Discord. We're now looking at Rocket.chat as a backup in case Discord goes down. But Slack is now completely out of the picture for our team.
I’ve advocated for an idea where Mattermost is to be used as a “bunker” where it is hosted on a raspberry Pi (or somewhere else) and acts as a digital bunker if your critical infrastructure (slack, teams, exchange?) is compromised somehow.
I know it's a tough spot, but if it were usable from GitLab with zero config that would be great for fallback.
To enable Mattermost, you can add the Mattermost external URL in the config file, and run `sudo gitlab-ctl reconfigure`. I'm wondering if that's something you've tried? https://docs.gitlab.com/omnibus/gitlab-mattermost/#getting-s...
* You get notifications for channels for which you have suppressed those notifications
* Some channels are marked as having new notifications, when they haven't
* Notifications for new messages in threads you are involved in are quite hard to find (horrible UX)
* Some UX choices are very confusing (you get a column of options related to notifications, and for some, the left option is the one leading to more notifications, for some the right option)
* There are some overlapping features that lead to inconsistent usage (channels vs. discussions vs. threads)
* Threads are hard to read, because follow-ups in threads are shown in a smaller font size. You cannot increase the font size at all in the desktop application
.... and so on.
Also, I tried to submit some bugs, but for that I'd need to have some information which only the admins have that run this instance, and in the end it was too much effort to get all that information together, so I didn't even bother.
I'm in a slack workspace that is constantly notifying me of a thread, but I can't make it read. Maybe it has something to do with the free messages limit. So the message is there, but cannot be accessed. Annoying as hell. I thought about submitting the bug to Slack, but then just let it go, and probably we'll just move to Signal or something.
Matrix is pretty straightforward on the server side of things, but the client UX is invariably mediocre. Vector—the official client—exemplifies everything that is wrong with Electron apps. Slow, clunky, poor UI, poor platform integration. With the default home server, it can take seconds for a message to go through. At least it’s far more customizable than Slack; it has an option for everything, which, as a power user, I quite like.
I haven’t tried Mattermost, but it looks like some of the important features aren’t FOSS, at which point it’s just another Slack as far as I’m concerned. I’ll gladly pay for support, but for SSO? Meh, might as well stick with Slack; at least everyone and their dog knows how to use it. (This is, of course, an opinion that stems partially from ignorance; I haven’t actually tried Mattermost, and if I do, I might fall in love with it. But my time is limited, and I can only evaluate so many products in a day.)
Not that Slack is much better here: their threading system has so many UI/UX issues. Ever had a thread with hundreds of messages? For your own sanity, I hope you haven’t. Ever tried to send an image to a thread from iOS? It’s possible, but only by pasting the image into the text field; the normal attachment button isn’t available, and Share buttons in other apps can’t send to threads. And, of course, the recent uptime issues.
That said, Element could certainly use less RAM, irrespective of Electron - and http://hydrogen.element.io is our project to experiment with minimum-footprint Matrix clients (it uses ~100x less RAM than Element).
Edit 1: I want to love it; the design is everything I could ever hope for in a chat platform. I even tried to contribute to Vector, but it was such a mess that I eventually gave up.
Edit 2:
> That said, Element could certainly use less RAM, irrespective of Electron - and http://hydrogen.element.io is our project to experiment with minimum-footprint Matrix clients (it uses ~100x less RAM than Element).
I'm not sure why this is a priority. Techies complain about RAM usage a lot, but if we have to choose between performance+power and a small memory footprint, we're going to choose the former almost every time. Take Telegram, for example: they have a bunch of native clients that perform amazingly well, although they do gobble RAM. Most of my technical friends use it as their primary social platform. It's not without issues, but it's really hard to go from something like Telegram Desktop or the Swift-based macOS Telegram client to Vector. And those clients aren't made by large teams--most (all?) first-party Telegram clients are each maintained by a single developer, if I'm not mistaken.
Does Element refer to the ecosystem as a whole, including EMS? The primary client? The core federation? It’s not obvious from a casual visit to element.io. I suppose if I said “Element web app,” that would be fairly clear, but I’m still in the habit of saying “Vector” from the days of Riot.
The protocol is still called Matrix.
Anyway, my point stands: Element Web/Desktop feels fairly unresponsive compared to something like Telegram Desktop. It looks so much nicer now, the UI layout is great, and it’s far more powerful than Telegram—yet, I can’t help but feel like I’m swimming through molasses even when dealing with moderately-sized groups. Try clicking around on different groups rapidly; you’ll likely find that you have to wait several seconds for the UI to update.
Wow – congrats!!
What have been the most important architectural decisions to achieve this?
XMPP is, well, extensible. If things don't match in the clients then that particular feature just doesn't work. These days all the clients pretty much try to match the feature set of Conversations. That applies to the servers as well. There is a server tester for that:
* https://compliance.conversations.im/
So the tooling is pretty good these days.
The great thing about XMPP is that basic messaging always works. That stuff is just too simple not to.
It even sent a 20 Mo MP4 like a champ, while Conversations sometimes chokes on not that high resolution photographs...
I run a small homeserver and use it to communicate with a group of about 20 friends. Most of them aren't "technical" people. We use it mostly for chatting and image/video sharing. We never use live calling (audio or video).
There have been a few bugs in the mobile apps, but for the most part, everything has been working fine.
The biggest issue is the UX. It's not as polished as the big players.
Like I said, it's close, I just don't think it's there yet.
The only thing that's a little cumbersome is requiring them to enter a custom server URL when the register/log in for the first time.
> The free Discord plan provides virtually all the core functionality of the platform with very few limitations. Free users get unlimited message history, screen sharing, unlimited server storage, up to eight users in a video call, and as many as 5,000 concurrent (i.e., online at the same time) users.
For a lot of small communities that aren't focused around commerce of any kind, Discord's free offering blows Element Matrix Services out the water. It's a non starter. If I could create a server with feature parity to Discord's free server, any new community I'd create I would definitely jump on EMS in a heartbeat, and I'd start trying to recreate communities currently within Discord, to be on EMS.
So like a very normal progression for Discord servers is that some niche sub-community wants to gather, and so they create a free server, and people join and there's all kinds of rich content that gets posted and curated and great discussions and then as it gets bigger, people running the community or people who want to support the community will boost the server with Discord Nitro for additional features like more slots for custom emojis (I can't communicate enough how important of a feature this is to Discord's success, even though it seems like minor window dressing).
That kind of model is what would justify a server starting to shell out money every month for EMS. I would note that Discord's pricing for this kind of level of community is tiered and not a per-user thing. You unlock more features based on how many users are paying for Nitro, going up a tier based on breakpoints of 2/15/30 Nitro Boosts per month. It doesn't cost more to have a tier 3 server if you gain more users. This is a big deal for fostering growth and unseating incumbent social networks (which is what Discord and Slack are).
Just some thoughts. I really want stuff like Element/Matrix to succeed!
and use it to communicate with a group of about 20
friends. Most of them aren't "technical" people.
I'm insanely curious about the human side of things here. How did you get them to buy into this idea in the first place? That sounds like quite an achievement.The non-technical folks in my life generally struggle with paths of least resistance (iMessage, etc) and it's hard to imagine getting them onto some alternative platform/protocol.
Previously, we were mostly using group texts or Snapchat/Instagram to communicate, so the biggest selling point was the fact that we can share full quality pictures and videos between iOS and Android people.
Not when it's down.
No hard data though. My mail server only ever went down when I upgraded the server and didn't check that everything was still working right away, or similar maintenance induced incidents. It never went down by itself.
Such systems only ever go down unpredictably on HW issues, or when overloaded/out of resources. Neither is very likely, because you're not trying to grow your service in any sense similar to VC backed enterprises. Most of the time it has constant very low load and resource use. And you can simply stop introducing changes to the system if you need more stability for some time. (stop updating, for example)
That doesn't happen for no reason. The vast majority of open source products I've used have terrible usability. I simply don't want to use them. I don't want to be beholden to corporations and walled gardens, but for me, the existing alternatives are worse in too many ways.
How often is Slack/Discord down? I mean it's not perfect, but I really honestly don't think I could match their uptime by self-hosting, as well as more on-call rotations for something that's not core product.
I very much prefer that for something that isn't core product, if it goes down I need to do exactly nothing for it to come back up, and that the engineers at Slack will be starting to work on it likely before I even realize it's down.
Hardware is quite reliable these days. And updates can be scheduled to be at a convenient time for the team.
If there are other people on the team that have _some_ technical skills then they can fix it..
IRC lacks quite a few features compared to other solutions, but the reduced complexity does bring very low operational complexity.
It's just not a business-friendly tool.
I'm an engineer and personally I'm fine with IRC, I'm just trying to be realistic here.
Also there are plenty of modern web clients for IRC, such as https://thelounge.chat/ or https://kiwiirc.com/ (which is supposed to work on mobile too).
Reality check: Most people don't use email programs anymore.
Also how do you get IRC to sync all conversation data, history, between your several desktops and phones, how do you send files, make calls, and thread conversations?
And you are starting to move goalposts here.. first it was uptime, then it was operations and now it's features...
And what about those web based IRC solutions? They are even easier to use than slack, have combined history, file sharing, etc.
Add to that the SaaS propaganda that hosting literally anything yourself is just too hard (it really isn't). Or this notion people are just too stupid to deal with anything more than the simplest possible web interface - Really? what do those people even do? Stare at Notepad all day? Of course not. They stare at various complicated software packages ranging from CAD, $spreadsheet abominations, SAP to various Adobe software packages. Sprinkle in a bit of hype for the latest new thing and presto.. </rant>
Guess you are not in enterprise.
If you are running networks and software on site, and they are business-critical, you have people and a plan for this. Or you don't, and suffer the consequences.
When self hosting you can get away with simpler systems that ends up being more stable and have higher up times for lower effort.
The reason you see cloud providers having issues is not because the thing is difficult, but because doing anything at huge scale ends up being difficult.
This is such a common misconception. The services I self-host was configured by me, if anything goes down (which they very rarely do), I know the exact cause and have it fixed in minutes. When some company's cloud service goes down I'm completely at their mercy. I also spend very little time on maintaining these services, just security updates, which are mostly automated.
Bottom line, maintaining and self-hosting services that has 1 or a few users is much less complex than services with millions of users. Hence, my uptime is better than Google's, Amazon's, and Azure's, etc.
The only thing "killed" XMPP was that proprietary made money, XMPP didn't.
Apart from that, it's alive and well. See Conversations for android, Prosody for server.
There might have been political reasons why google dropped XMPP, but it would also make sense as a purely technical decision.
What makes you think so? If Conversations was draining my battery, I would have noticed by now, I'm pretty sure that Facebook Messenger is worse in this aspect...
Has this been fixed in XMPP?
This is plain untrue. Yes it was invented a long time ago, but thanks to the extensibility it has evolved over time just as the way people use it has changed. This evolution is a healthy and necessary part of an open ecosystem.
I know it frustrates people that modern features don't work in stagnated clients such as Pidgin and Adium, but modern clients support all the things you would expect.
Servers and mobile clients have supported mobile-friendly traffic and connection optimisations for many many years now.
> There might have been political reasons why google dropped XMPP, but it would also make sense as a purely technical decision.
Google contributed extensions to XMPP, the same way they contribute to other internet standards. I think they were quite comfortable with this. The XMPP-based Google Talk was their longest-running messaging solution after all...
That's true for email also. So that's not an evidence for anything lacking in xmpp. The plain fact is that Google killed it due to their greed.
Scrolling back is atrociously slow, and it doesn't even seem to have a search feature !
I'd take XMPP alternatives like Conversations, Jitsi, Pidgin any day ! (And Element of course.)
Element (the primary Matrix software) definitely has Slack and Discord in its sights.
I don't think there are any serious "self-hosted Slack-like" contenders that are XMPP-based right now. You can piece components together (yay, standards!) and I did exactly this for the IETF's XMPP deployment recently. But it's far from being a cohesive easy-to-deploy product. Simply because nobody is building that right now. It takes time and resources and there's no money in it.[1]
People who do set out to build Slack clones (projects like Mattermost and Rocket Chat) and earn money don't have features such as federation on their priority list and don't build on top of Matrix/XMPP. They roll their own custom protocols and as far as I can see they are fairly content with that decision.
[1] There's even less money it, but nevertheless I am currently working on such a self-hostable "package" for XMPP. However rather than focusing on the team chat use-case (Slack/etc.) I'm focusing on personal messaging (WhatsApp/etc.): https://snikket.org/ if you're interested. It's possible I will broaden the scope one day.
EDIT: typo
The essential problem IMO is how to replace SMTP. No one has proposed and implemented an alternative, to my knowledge. So I decided to[1]. The current draft omits federation (although I wouldn't rule it out in all cases yet).
[1] https://github.com/networkimprov/mnm/blob/master/Protocol.md
And you can keep using all the normal MUA's on desktop and mobile.
Setting up a new server isn't easy unless you hire an outside service provider, and if you're willing to do that, Slack et al offer a nicer UX than the well known email/webmail clients.
Orgs with sufficient IT resources commonly do run internal SMTP servers.
> problem IMO is how to replace SMTP.
Sadly SMTP is probably one of the parts of Mail which have aged best. Enforcing the usage of some (currently by design optional) features wrt. authentication and similar at the cost of backwards compatibility and you have all you need from the delivery protocol.
BUT:
- IMAP and similar is much worse.
- Mail bodies are a big mess it's always fascinating for me that mail interoperability works at all in practice (again you can clean it up a lot, theoretically, but backwards compatibility would be gone).
- DMARC, DKIM and SPIF which handle mail authenticity have a lot of rough corners and again for backward compatibility are optional. Again it's not to hard to improve on but would brake backwards compatibility.
The main reason mail still matters is because it's backwards compatibility, not just with older software but also with new software still using old patterns because of the (relative to the gain) insane amount of work you need to put into all kinds of mail related components. But then exactly that backwards compatibility is what.
(Yes, I have read the "Why TMTP?" link and I have written software for many parts around mail including SMTP, and mail encoding. The idea that SMTP is at the root of the problem seems to me very strange. Especially given that like I mentioned literally every other part of mail is worse then SMTP by multiple degrees...)
EDIT: Just to prevent misunderstandings one core feature of mail is the separation of mail delivery and mail authenticity, in the sense that you don't need the mailman to prove the authenticity of a mail. At most the legal/correct/authentic delivery.
The opposite is also true.
TMTP also covers most IMAP/POP use cases. And it allows short, plain-text messages (see Ping) to make first contact with others -- necessary when that server has less restrictive membership requirements.
Authenticity is a double-edged sword. For certain confidential content, you want the recipient to know that it originated with the sender, but you don't want anyone else to know that in the event the content is leaked or stolen.
I believe the extinction of email for person-to-person & app-to-person correspondence is a foregone conclusion, due principally to phishing. The question is what should we do now, and the answer is clearly not chatrooms (which are of course useful in certain circumstances).
One email like protocol that properly handles this is NNTP.
Protocol: https://github.com/networkimprov/mnm/blob/master/Protocol.md
Why TMTP? https://mnmnotmail.org/rationale.html
Follow: https://twitter.com/mnmnotmail
Regarding the actual modes of discussion I was thinking of though, usenet and email are mostly the same.
For the most part, they are and many readers support both protocols (or at least they did in the past). The nice thing about NNTP is that it doesn't require maintaining a separate archive or having someone send you an mbox file to import. Just subscribing to the appropriate groups was sufficient (depending on the article retention policy).
The same could be accomplished with email if you only allow connections to the SMTP and IMAP server from within the corporate network. That is, nothing external can connect to those servers, which is fine if it's only used for internal communication.
Ahem, I'm trying to bury email, not save it -- not unlike Slack :-)
Sorry my dude, but business runs on email. Saying lets get rid of it is as naïve as saying lets get rid of Excel. It's just not going to happen.
1) Set up a second TMTP service where customers and/or suppliers can create accounts, along with employees who need to interact with them.
2) Have some employees join a third-party service which is open to all involved in your field. There is a risk of phishing in this case, but:
. a) anyone you haven't previously contacted is limited to short, plain-text communications to you (see Ping in protocol), and
. b) such third-party services would typically charge a fee to members and impose a small cost per-ping, and
. c) you know that you're dealing with unknown entities with possibly malicious intent.
TMTP clients support active logins to accounts on any number of TMTP servers, just as browsers support multiple active connections to websites.
But I signed up and will keep it on my radar as it matures.
1. There is virtually zero user-facing documentation. Need to know how to backup keys, verify another user, or what E2EE means? Ask your server operator. Basically the onus is on operators to document this stuff for their users. Except the stuff we're documenting is hard even for server operators, and especially challenging to document in a way that both nontechnical and technical users can understand.
2. Because this stuff is challenging even for more technically minded users to understand, it leads to a kind of burnout for interested non-technical users where they learn all they can about some feature and how it works at a high level from out of date random blogs, try to use the (complex, multi-step) feature, but then something won't work, and it isn't be clear whether it was because the user did something wrong or because the clients or server implementations are broken
3. Issues where core functionality is broken (e.g. two mutually verified users on my homeserver haven't been able to talk to each other in months -- see [1], [2], [3]) languish for months with zero response from maintainers.
4. While core functionality is both broken and undocumented, the maintainers announce rabbit hole features that no one asked for and seem very much like distractions, like their recently-announced microblogging view/client[4]
In short the Element maintainers have shown little interest in making the platform accessible to the people who need its differentiating features the most, and have prioritized the "mad science"/technical aspect of their platform at the expense of the human element (end-users and operators).
It'd be cool if Element used their resources to hire some UX folks and community advocates whose sole focus is addressing the horrid accessibility of their platform. I think most users would rather see that than further "mad science".
[1] https://github.com/vector-im/element-ios/issues/3762
[2] https://github.com/vector-im/element-ios/issues/3572
Also, it seems the FAQ answers several of your points: https://element.io/help
That FAQ is a great start, but it's not sufficient for non-technical users. It's not easily searchable, it doesn't provide screenshots, and it doesn't go into enough detail for each item (e.g. describing what can go wrong + troubleshooting).
Thanks for pointing it out though
it also doesn't take gigs of memory on client devices.
It's bad enough team comms go over Slack so much now, at least we have email fallback. What scares me is for the teams that use Slack for system alerting.
At one of my old jobs, we had SMS via two physical/hardware devices in our data center. One had a Telstra SIM card and the other had an Optus SIM card. (They were plugged into the same machine, but we had plans to put a second one in another data center before I left).
If you really care about alerts, you should have physical hardware doing your SMS messages via two different point-of-presences.
And I'm in that boat of depending on Slack for alerting... in fact my team was also waiting over the holidays to deploy more robust non-Slack-based alerting (in our defense the product is only a few months old and only now starting to scale to any real volume).
I used to work for a place that had a FY that ended in summer. We had a lot less problems with stuff being shoveled out the door at Thanksgiving and Christmas because nobody was trying to finish their year-end performance goals over the Holidays.
I think what I'm implying is that management creates this issue, but we are complicit.
The holidays are actually the perfect time for Slack to roll out a risky deployment, as it has to be their lowest usage time. So it would make sense if something was pushed out last week or the week before. And everything probably seemed fine.
And then this morning they suddenly realize this new feature does not perform under load. And to make matters worse, the new feature has been out long enough to make any sort of rollback very tricky, if not impossible. Which means they'd need engineers to desperately hack out, test and deploy a code fix.
If this is the scenario, I do not envy them at all.
It broke, which was a major problem, this meant that senior management were being phoned (ho), and relatively high middle managers were on site to deal with the fall out. Of course most suppliers were also closed so everything was harder to fix.
There's good reasons not to do changes when places are closed, or at least skeletoned, for 2 weeks.
Of course 2020 saw an increase, but it was smeared over a week or a month rather than being a big jump in a single day after everyone's holidays.
It seems a bit presumptuous to assume that's at fault here, given their age.
Hopefully email won't be your backup. I've seen that done. Alerts get filtered and ignored, often by accident.
I'll take this moment to remind everyone of their human tendency to read meaning into random events. There's no evidence to suggest New Year traffic has caused this, and outages like this can happen in spite of professional and competent preparation.
Hugops for their team, I hope they get it back soon.
On the one hand, sure we don't specifically know what's going on. On the other hand, it's the first Monday in the new year and they went down shortly after the start of the business day Eastern time; it could be coincidence, but it would be a remarkable coincidence.
It could be anything really- my post was more about how situations like this can happen to even the most prepared. The assumption it has something to do with NY tends to assume very trivial, silly mistakes. Especially with no information, that seems a bit uncharitable.
I'm posting this because I found a lot of people don't know that Zoom includes a complete chat client that includes channels.
And #HugOps to the engineers at Slack working on this. I appreciate that they posted a periodic update even when there was no news to report: "There are no changes to report as of yet. We're still all hands on deck and continuing to dig in on our side. We'll continue to share updates every 30 minutes until the incident has been downgraded."
https://www.washingtonpost.com/technology/2020/12/18/zoom-he...
This was an executive, not just an employee. That's a huge distinction and I can't help but think you intentionally downgraded his position to cover-up his behavior. "Just an employee" "Not a big deal"
But when you read the allegations, they seem like a very big deal that an executive was spying on users, giving their information to the Chinese government explicitly for oppressive purposes, including folks who are not in China, and went out of his way to personally censor non-Chinese groups meeting to discuss the Massacre-Which-Cannot-Be-Mentioned.
I would say the headline understates the gravity (it's very much a 'by-the-books' headline that you KNOW went through ten levels of Legal), and that your hand waving here feels much more dishonest than the headline.
> and other employees have been placed on administrative leave until the investigation is complete.
Zoom at least suspects he did not act alone.
And it's also undeniable that the consequences for Zoom (really, just needing to fire a few people, and not even the people who designed those controls if there were any) are so minimal that they have no incentive to strengthen those controls.
For some organizations (mine included) the benefits of Zoom outweigh the risks of Zoom having proven itself to not have those controls, namely the possibility of both political and corporate espionage. As with all things, YMMV.
America is fucked up, that doesn't mean that other countries aren't also fucked up or aren't doing worse things with the data they collect.
Essentially, the China case proved Zoom is willing to cooperate with a nation state. The US is the nation state we live in, Zoom is HQ'd here. Therefore, the risk to us is high.
As an aside, the organ harvesting idea comes from the Fulan Gong, who are similar to Chinese Scientologists. It is not clear to me that their claims are accurate.
No one imprisoned in Guantanamo Bay is a US Citizen and neither is Assange.
The US prison system is super fucked up but it is not the same as ethnic cleansing.
You are comparing apples to concentration camps.
https://www.channel4.com/news/factcheck/factcheck-were-mass-...
https://www.snopes.com/ap/2020/09/18/more-migrant-women-say-...
And 70% of those people in those camps are released within 30 days, often times within one week back to their country of origin (or given asylum).
Are the "brown people" in camps along the border a single, ethnic minority? Are all "brown people" in the country subject to arrest and under surveillance just for being "brown"?
I think you should I know I -- and probably others, are reading this as "b-b-but, they're not US Citizens, so they don't deserve [the same] rights"
I hope that's not what you mean, because if it is, that's really fucked up.
That's a glib retort.
A takeaway from your position is that it's ok so long as you do it to citizens of other countries.
> it is not the same as ethnic cleansing.
See the above.
That's always been the difference between the US and China and why so many countries have hatred for us and yet little to none for China. They don't fuck with other countries on the level that we do.
The fear on this forum is imagined political thriller more than realistic.
Every technologist is grifting off the military industrial complex.
I'm not condoning Zoom's actions but this is hardly a problem unique to Zoom. Few if any businesses will stand up for consumers and citizens unless it's directly aligned with their profit motive. In this case, the business choice is to operate or not in mainland China. If they choose to stand up against the Chinese government they're going to have difficulty continuing to operate in China and risk losing that entire market.
Google played this PR game many years ago in China (rejecting some of the governmental policies) and ultimately caved to Chinese policies to do business there.
Businesses are not the organizations we should look to for empowering people, that's simply not their goal no matter how much their marketing team may want to sell that idea by following trending (popular) social movements that they've already done market studies on to assess potential fallback.
What other business in this space has given China unfettered access to US users and data? I'm not aware of it occurring with Webex, Teams or go2meeting. The "one rogue employee" thing falls flat pretty quickly when they're the only ones that had this issue.
This feels like their encryption thing all over again, there's an "oversight" that is equivalent to a backdoor that only gets fixed when they get caught.
The fact is all of the US businesses operating in China give surveillance ability to the Chinese government for the Chinese users and are operating in an ethically questionable space being primarily based outside of China, at least in my opinion.
It's really not too different than the businesses sharing US citizen data to the US government, much of which Snowden and others before him exposed. I suspect there's a lot more surveillance going on everywhere than the general public know about and the businesses best positioned to do the surveillance are probably doing it.
I'm amazed they use Slack at all. Let alone as a fallback.
https://about.gitlab.com/blog/2015/08/18/gitlab-loves-matter...
> Like many companies in the last year we've switched to using Slack to improve internal communication. [...] Since Slack doesn't offer an on-premises version, we searched for other options. We found Mattermost to be the leading open source Slack-alternative and suggested a collaboration to the Mattermost team.
I'm not really sure why it's 'GitLab Mattermost' and not (at your link) 'GitLab Nginx' et al. though.
(We use GitLab and Mattermost (integration) where I work. I've been 'remote / WFH' for the past 7 years.)
They use a ton of services.
Likely you don't want your backup to be one of your systems and another part of the company probably uses Zoom already so it is probably easy to fail over to that.
This includes many proprietary ones, we generally choose the product that will work best for us, considering the benefits of open source, but not excluding proprietary software.
Mattermost is not part of the single application that GitLab is. There is a good integration between with GitLab and our Omnibus installer allows you to easily install it. But it is a separate application from a separate company.
That is one reason we did not go with Mattermost.
This isn't a slam at devops, it's about the need for institutional information hiding; not everyone needs to know about and weigh in on every decision being made.
https://www.nbcnews.com/better/business/slack-updates-privac...
This is the reason these "closed" ecosystem apps like Slack/Zoom are multi billion dollar companies and have massive uptake. Simple and easy to user.
--
Though, I do agree wholeheartedly with your sentiment that the Slack team needs all the positive vibes they can get right now.
However lot of the monitoring which alerts on slack and other automatic notifications are critical for many teams.
How would people react? What would engineers do to recover? I always found that idea fascinating.
Imagine Google saying tomorrow that they lost all the accounts and emails. What kind of impact the world will have?
Slack is mostly real time communication, at least for me. There are a few bits and bobs that really should be documented that are in the messages though.
You not only have backups in place, you have documentation in place, including a back-up vendor who has copies of the documentation and can staff up workers to get it up and running again without any help from existing staff.
And we tested those scenarios. I'm not sure which dry runs were less fun - when you got paged at 3 AM to go to the DR site and restore the entire infrastructure from scratch... or when you got paged at 3 AM and were instructed to stay home and not communicate with anyone for 24 hours to prove it can be done with out you. (OK, so staying home was definitely more fun, but disturbing.)
How many companies really plan for an event where their entire infrastructure goes offline and their entire team gets killed? Does even companies like Google plan for this kind of event?
But based on my experience, the initial recovery planning is the hard part. The documentation to tell a new team how to do it isn't so painful once the base plan exists, although you do need to think ahead to make sure somebody at your back-up vendor has an account with enough access to set up all the other accounts that will need to be created, including authorization to spend money to make it happen.
In some ways losing your ERP and it's backups would be harder to recover from than both sites burning down, insurance would cover that at least.
It's hearsay, but I was once told that achieving "black start" capability was a program that took many years and about a billion dollars. But they (probably) have it now.
It's most often referred to in the electricity sector, where bringing power up after a major regional blackout (think 2003 NE blackout) is extremely nontrivial, since the normal steps to turn on a power plant usually requires power: for example, operating valves in a hydro plant or blowers in a coal/gas/oil plant, synchronizing your generation with grid frequency, having something to consume the power; even operating the relays and circuit breakers to connect to the grid may require grid power.
The idea here is presumably that Google services have so many mutual dependencies that if everything were to go down, restarting would be nontrivial because every service would be blocked on starting up due to some other service not being available.
Since 9/11, more than you might think. For example Empire Blue Cross Blue Shield [1] had its HQ in the WTC.
https://www.computerworld.com/article/2585046/empire-blue-[1... cross-it-group-undaunted-by-wtc-attack--anthrax-scare.html
And what a blast from the past:
> Some of the temporary locations, such as the W Hotel, required significant upgrades to their network infrastructure, Klepper said. "We're running a Gigabit Ethernet now here in the W Hotel,'' Klepper said, with a network connected to four T1 (1.54M bit/sec) circuits. That network supports the code development for a Web-based interface to the company's systems, which Klepper called "critical" to Empire's efforts to serve its customers. Despite the lost time and the lost code in the collapse of the World Trade Center towers, Klepper said, "we're going to get this done by the end of the year."
> Shevin Conway, Empire's chief technology officer, said that while the company lost about "10 days' worth" of source code, the entire object-oriented executable code survived, as it had been electronically transferred to the Staten Island data center.
It's pretty much part of the basic day-to-day life in some industries.
I sat in on a DR test where the moment one of the Auckland based ops team tried asking the Wellington lead, the boss stepped in and said "Wellington has been levelled by an earthquake. Everyone is dead or trying to get back to their family. They will not be helping you during the exercise."
https://www.nytimes.com/2014/11/19/magazine/the-secret-life-...
They had backups and were able to recover data and systems.
By the time I got there, they were somewhat functional.
The biggest problems were the lack of knowledgeable personnel, not lost data or systems.
I was there as a consultant and didn't know anyone there when I went.
I won't provide any details out of respect for those fine people, but the grief was so thick, you could have cut it with a knife. As I said, I didn't know anyone who was there (or wasn't there) but after a day, I wanted to cry.
The rest of the world may not be so energetic re: their accounts and data, so it would be painful for many, it depends on their how much risk they are willing to experience.
Being in DR, it is very difficult for businesses to allocate the time and resources to good planning - for many, DR is an insurance policy. Staff: engineering and development are focused on putting out fires however, a real DR is more than most companies can handle if they have not planned accordingly or practiced through testing failover/normalization processes as well as performing component-level testing.
HN comments at the time: https://news.ycombinator.com/item?id=487497
The site relaunched a month later and shut down for good a year after that
IME messages just fail to send with Slack, then you can retry but they're not properly idempotent and you end up sending the messages twice.
It's really poor.
https://battlepenguin.com/tech/matrix-one-chat-protocol-to-r...
It works fairly well.
Really a huge part of IRC's difficulty and beauty is in not having a markup language, but most of that beauty is for the eyes of the developer, not the user.
I like the concept of Matrix. That's kind of what they're trying to do by creating an open protocol, but when I looked at implementing a client it was non-trivial. For IRC, you can usually send someone a telnet log of you joining an IRC server and they could implement a client. I don't get the impression that that's true for Matrix.
Only centralized history/logging and search would need to be bolted on if needed. In the non-centralized case your IRC client takes care of all of that.
Since I would see Slack more of a replacement for phone calls or hallway discussions. Neither of which typically has any logs or recordings (and I wouldn't want to work somewhere that did keep such logs).
I'm sure you know this already, but that status page isn't worth the cycles on your CPU, you would be better served asking the toaster if AWS is functioning properly than checking that status page.
[edit] Nevermind, I just needed the right combination of terms to find it: https://docs.aws.amazon.com/ram/latest/userguide/working-wit...
The dashboards are all green. (which doesn't mean that much ... I'm aware...)
* iMessage, which likely handles something in the range of 750M-1B monthly actives.
* WhatsApp, 2B users [1], though no clarity on "active" users.
* Telegram, 400M monthly actives [2]
* Discord, 100M monthly actives [3]
* Slack, 12M daily actives [4]
* Teams, which is certainly more popular than Slack, but I shudder to list it because its stability may actually be worse.
The old piece of wisdom that "real-time chat is hard" is something I've always taken at face-value as being true, because it is hard, but some of the most stable, highest scale services I've ever interfaced with are chat services. iMessage NEVER goes down. I have to conclude that Slack's unacceptable instability, even relative to more static services like Jira, is less the product of the difficulty of their product domain, and moreso something far deeper and more unfixable.
I would not assume that this will improve after they are fully integrated with Salesforce. If your company is on Slack, its time to investigate an alternative, and I'm fearful of the fact that there are very few strong ones in the enterprise world.
[1] https://blog.whatsapp.com/two-billion-users-connecting-the-w...
[2] https://techcrunch.com/2020/04/24/telegram-hits-400-million-...
[3] https://wersm.com/discord-reaches-100m-monthly-active-users-...
[4] https://www.cnbc.com/2019/10/10/slack-says-it-crossed-12-mil... (this was also announced on Slack's blog, but that's down).
The really impressive thing about Discord's scale is the size of their subscriber pools in the pub-sub model. Discord is slightly different than Slack in the sense that every User on a Server receives every message from every Channel; you don't opt-in to Channels as in Slack, and you can't opt-out (though some channels can be restricted to only certain roles within the Server, this is the minority of Channels).
Some of the largest Discord servers have over 1 million ONLINE users actively receiving messages; this is mostly the official servers for major games, like Fortnite, Minecraft, and League of Legends.
In other words, while the MAU/DAU counts may be within the same order of magnitude, Discord's DAUs are more centralized into larger servers, and also tend to be members of more servers than an average Slack DAU. Its a far harder problem.
The chat rooms are oftentimes unusable, but most of these users only lurk. Nonetheless, think about that scale for a second; when a user sends a message, it is delivered (very quickly!) to a million people. That's insane. Then combine that with insanely good, low latency audio, and best-in-class stability; Discord is a very impressive product, possibly one of the most impressive, and does not get nearly enough credit for what they've accomplished.
For comparison; a "Team" in Microsoft Teams (roughly equivalent to a Discord Server or Slack Workspace) is still limited to 5,000 people.
Keep in mind you're comparing daily active users vs monthly active users. I'd guess most slack users are online weekday for pretty much the entire day (because it's for work and your boss expects you to be online), whereas a good chunk of discord users are only logging in a few hours a week when they're gaming.
Minecraft official server: 190k online users. | Fortnite official server: 180k online users. | Valorant official server: 170k online users. | Jet's Dream World (community): 130k online users. | CallMeCarson server (YouTuber): 100k online users. | Call of Duty official server: 90k online users. | Rust (the game) official discord: 80k online users. | League of Legends official server: 60k online users. | Among Us official server: 50k online users.
Their scale is insane. Even with their usage spiking during after-hours gaming in major countries, their baseline usage at every hour of the day, globally, makes it one of the most used web services ever created.
Slack's DAU and MAU numbers are probably pretty close to one-another. Discord's MAU/DAU ratio is probably bigger than Slack's. That just means that Discord is, again, solving a harder problem; they have much bigger (and more unpredictable) spikes in usage than Slack. Yet, its a far more stable and pleasant product.
Well for the real time side, I can't tell you how big a boon it's been to build our platform on top of Elixir/BEAM. Hands down the best runtime / VM for the job - and a big big secret to our success. Where we couldn't get BEAM fast enough - we lean on rust and embed it into the VM via NIFs.
2021 is the year of rust - with the async ecosystem continuing to mature (tokio 1.0 release) we will be investing heavily in moving a lot of our workloads from Python to Rust - and using Rust in more places, for example, as backend data services that sit in front of our DBs. We have already piloted this last year for our messages data store and have implemented such things as concurrency throttles and query coalescing to keep the upstream data layer stable. It has helped tremendously but we still have a lot of work to do!
To help scale those super large servers, in 2020 we invested heavily in making sure our distributed system can handle the load.
Did you know that all those mega servers you listed run within our distribution on the same hardware and clusters as every other discord server - with no special tenancy within our distribution. The largest servers are scheduled amongst the smallest servers and don't get any special treatment. As a server grows - it of course is able to consume a larger share of resources within our distribution - and automatically transitions to a mode built for large servers (we call this "relays" internally.) At any hour, over a hundred million BEAM processes are concurrently scheduled within our distributed system. Each with specific jobs within their respective clusters. A process may run your presence, websocket connection, session on discord, voice chat server, go live stream, your 1:1/group DM call, etc. We schedule/reschedule/terminate processes at a rate of a few hundred thousand per minute. We are able to scale by adding more nodes to each cluster - and processes are live migrated to the new nodes. This is an operation we perform regularly - and actually is how we deploy updates to our real time system.
I was responsible for building and architecting much of these systems. It's been super cool to work on - and - it's cool to see people acknowledge the scale we now run at! Thank you!! It's been a wild ride haha.
As for scale, our last public number perhaps comparable to Slack is ~650 billion messages sent in 2020, and a few trillion minutes of voice/video chat activity. However given the crazy growth that has happened last year due to COVID - the daily message send volumes are well over the 2 billion/day average.
I think the big things that prevent it from being adopted more for professional use is the lack of a threading model (even though I hate it when people use threads in Slack) and the whole everyone in every channel except for role-based privacy settings. The second one especially is a big deal because you can't do things like team-only channels without a prohibitive amount of overhead.
That said (with zero knowledge of their architecture) I have to feel like both of those missing features aren't too terribly hard to build. Its very likely Discord is growing as a business fast enough on the gaming and community spaces they don't feel the added overhead of expanding into enterprise (read: support, SLAs, SOC, etc) makes sense and are waiting until they need a boost to play that card.
They do have a threading model now (if you are talking about replying to a message in a channel and having your reply clearly show what you are responding to). If you are talking about 1-on-1 chats with other people in your same server then yes, that is still lacking IMHO in discord. The whole "you have to be friends" to start a chat (or maybe that's just for a on-the-fly group) is annoying.
We have a few bots we've integrated with things (deployment, stats, etc).
We use it for all our voice/video calls.
Edit: We've got roles setup well for things like contractors, devs, marketing, etc, so it's easy to lock down different conversations in channels.
It's been fantastic.
The only thing I'm not a huge fan of is the (IMO) poor implementation of threaded discussions.
Edit: it definitely has issues with connectivity from time-to-time too, but not bad overall.
TBH, I'm not sure why companies use Slack (I use it for other organizations, so have experience with it too, but not extensive).
ITT: Anecdotes
That being said, individual instances of the app are notoriously unstable causing random annoyances. But, I am on a very early build of Teams, which is buggy by definition.
edit: oh, sorry, i did say 'status page' in the first part. But I kinda meant update % tracker like the parent.
It's listed as "Uptime for the current quarter"; if they mean that as "calendar quarter", i.e. since the start of the year, then we aren't even 100 hours into the quarter so we should be well below 100% by now.
A status page that doesn't get updated during an outage is about as much use as a solar-powered flashlight (without built in power storage).
It's like when people push the elevator button repeatedly if it's taking a while to arrive, only pushing the elevator button doesn't cause it to take even longer.
Probably not a priority though.
(edit: to clarify: not affiliated in any way, just a fan)
That said, we're pushing up against the limits of our free plan with Slack and will likely deploy a matrix server in due course.
For example, Mirage is Python/QT and quite fast in my experience. There are Rust clients, C++ clients, terminal based ones, etc.
I'd like to get off of Element desktop/web for a couple of reasons, but I need those features. I'd help implement them myself, but that's beyond my skill level.
Edit: For anyone else wondering, matrix-commander [0] looks like it may be workable if a cli tool is acceptable for your usecase.
[0] https://matrix.org/docs/projects/client/matrix-commander/
I'm planning on looking through the GUI ones at some point, but don't have time now.
Note: I wanted to try it out for a while, but haven't yet.
looks through Inbox of 850 new aws, batch job and logging messages
oh yea, that's right..
"Hey, our site has been down for 2 hours, why aren't you doing anything"
Looks at 850 unread messages in ops-notifications folder
ooh yeah, that's right..
In my organization it's spelt "deleted items"
Oh wait, how would you share the link...
I lost the login for our shared AWS account. Mind sending it to me here?
* work on code
* update JIRA
* complete required trainings
* work on my peer reviews (Workday/Okta are up)
* review tech specs
Deploying is actually a very small part of my job.There's also this perverse incentive to Slack all the things. Lots of CI notifications are sent through it. Some org processes are implemented as workflows. There's been talk of how wonderful it would be to hook up tasking and work tracking to slash commands. I and others often use Slack instead of the 'official' tool to video call each other.
An outage like this is still really disruptive. It's not like everyone realizes what's going on immediately or at the same time; we have backup tools, but our turn radius is pretty wide. Some of us can't even communicate effectively without memes, too, and backup tools don't have a giphy integration.
EDIT: Do your CI integrations fail if Slack can't be contacted? Do those failures fail your pipeline? Whoops!
However, I drive everything through slack - GitHub, linear, calendars, Notion, support emails, etc. I have notifications turned off for every service we use except for slack. This allows me to effectively ignore everything except for slack. These types of outages destroy that workflow for me.
But it was roughly "one large impact a month, for six months", with large caveats that upper management for whatever company had to be working with the product during that month.
Large companies don't care if X service went out during the night and impacted someone not in their timezone.
If the CTO notices that he can't use something with the same regularity that he gets paid, then it doesn't take long for it to stick in their mind. But migrating everything is _so painful_ that the majority of large companies will do anything they can to avoid moving away.
This is a key point is the popularity amongst VCs in investing in B2B SaaS. I take their (and your) word for it. But honestly, I don't actually understand this.
Why is migration so hard?
This is not to mention the fact that half our staff aren't hugely technical, so have actively _learnt_ how to use Slack and it's features around notification control (things that may come "naturally" to the tech-savvy crowd on HN), @-things, bots, etc, and they would need to re-learn a new tool that is going to work in a different way.
This would be a substantial effort for us, and we're a small company. Are there ways to materially minimise this cost?
Why would anyone make it easy
Things that are easy to migrate from get replaced by things that are hard to migrate from, eventually.
IRC is incredibly easy to migrate from.
The parent meant a law as in "a law of physics", not a piece of legislation.
Plus, it will just take a long time to get everyone on board and using the replacement system. My department is slowly plodding towards using Teams over Slack, but there are enough hold-outs (my sub-department being one of them) that it still doesn't have wide-spread adoption.
If you had a small shop with a dozen tech-savvy people and Slack became a problem which was used exclusively for quick business chats, you could probably push a change to another chat platform the next day. You might struggle when you have thousands of employees, some that needed training to use Slack and still aren't that proficient.
* Getting out of the Enterprise Contract, or waiting for the year to end. * Training people on new software. * Loss of productivity. (1) Learning a new UI, processes, workflows -- both individually and organizationally. A feature or concept in "Tool A" may exist in a completely different form in "Tool B". Or not exist, and then people need to adapt to and work around the missing feature. (2) Missing out on needed information due to the above. Ultimately, software exists to move and transform data, and when you change the software people have to adjust. Sometimes that doesn't go great. "Oh, I didn't realize I needed to check this checkbox".
Another way to say this is "organizational inertia", which is a fancy term that means "it's hard for people to adjust to change".
And you might think developers and other technical people would have an easier time of it. They (we) do, but not to the extent you may expect. I've been on the front lines of a handful of migrations that affected only the IT staff, and it was a long and arduous process each time.
Man it bothers me so much when applications change their UIs on updates for no apparent reason other than "it looks better".
IntelliJ changed the way build and debug buttons looked in some update and it took me days to get used to it and I could find them in a snap again. Slack did a couple of no-reason changes as well.
The really big one, for companies of a certain size / cash flow, is compliance. Companies spend a lot of time developing compliant work flows around a service like Slack.
Migrating to another service requires rewriting the compliance narrative. The current compliance people might not have the confidence or willpower to do that effectively, and can raise legal objections to any such migration indefinitely.
I think we should push for a metric where "up" means 100% of people that want to use the service are able to use the service. If 1% of users can't send messages, then that should count as a full-blown outage and should start counting against whatever SLA they advertise.
The underlying problem here is that apparently everyone lies about uptime, so if you don't, that looks bad to potential customers. I fear that we will have to push for some legal regulation if we want accurate data, and ... people will probably be opposed to that.
I was wondering why the link from a Jira wasn't opening in slack, the page eventually timed out and gave me a link to status.slack.com where it told me everything was peachy. Cue me wasting time trying it again because apparently there was no issue with slack..
Honestly, I could live with a 99.50% SLA, if that's what it really was. After today's probably full-day outage, they'd just have to be extra careful for the rest of the year (or pay me money). Kind of sucks when it's 1/4 that you blow your year's SLA budget though.
I mean, that’s nice to say, but how do you measure/prove it?
Certainly, having the SLAed party check themselves is silly. But what are the other options? If it was up to the customer, customers could make up faults to get free service. (Since it’d be up to the customer to prove, and customers are generally less technical than vendors, you’d have to expect/accept very non-technical — and thus non-evidentiary! — forms of “proof”, e.g. “I dunno, we weren’t able to reach it today.” Things that could have just as well been their own ISP, or even operator error on their side.)
IMHO, contractual SLAs should be based on the checks of some agreed-upon neutral-third-party auditor (e.g. any of the many status/uptime monitoring services.) If the third party says the service is up, it’s up in SLA terms; if the third party says the service is down, it’s down in SLA terms.
(And, of course, if the third party themselves go down, or experience connectivity issues that cause them to see false correlated failures among many services, that should be explicitly written into the SLA as a condition where the customer isn’t going to get a remedial award against the SLA, even if the SLAed service does go down during that time. If the Internet backbone falls over, that’s the equivalent of what insurance providers call an “act of God.”)
But in a neutral-third-party observer setup, you aren’t going to get 100% coverage for customer-seen problems. An uptime service isn’t going to see the service the way every single customer does. Only the way one particular customer would. So it’s not going to notice these spurious some-customers-see-it-some-don’t faults.
So, again: what kind of input would feed this hypothetical “100% of customers are being served successfully” metric?
ETA: maybe you could get closer to this ideal by ensuring that the monitoring service 1. is effectively running a full integration test suite, not just hitting trivial APIs; and 2. if gradual-rollout experiments ala “hash the user’s ID to land them in an experiment hash-ring position, and assign feature flags to sections of the hash ring” are in use by the SLAed service, then the monitoring service should be given N different “probe users” that together cover the complete hash-ring of possible generated-feature-flag combinations. Or given special keys that get randomly assigned a different combination of feature-flags every time they’re used.
Google published a paper last year describing this approach to measuring uptime: https://blog.acolyer.org/2020/02/26/meaningful-availability/
The idea is to define availability as "the probability that the site 'appeared' to be down for a random user, averaged over a time window of size w". You can choose a particular value of w and look at trends over time, or you can plot availability as a function of w to understand patterns of downtime.
If you look at the history page you can see its not 100% for every month: https://status.slack.com/calendar
Without being logged in, things are as fast as they usually are -- but post log-in, SLOWWWER THAN MOLASSESS...
I tried this several times; why this is, I can only wonder...
To quote Bill and Ted... "Strange things are afoot at the Circle-K..."
I don't think that's true. For example, if you hide a thread while logged in, it remains hidden when you return.
But also, the status page still proudly proclaims that the "Uptime for the current quarter: 100%" — which is clearly false at this point.
Slack, notion and AWS all at same time seems unlikely
Quite glad we never moved any critical ops work into Slack bots, since we don't control Slack.
I wonder how many companies (like mine) have literally ground to a halt because of this? Do other companies have a risk-documented backup plan B for times like this? Presumably the default is for everyone to resort to email?
More worryingly is the number of ChatOps processes and alerting/observability systems that are in place around Slack.
Not being able to chat with co-workers for an hour or two is fine, but not being able to safely manage CI/CD/deployments is a big risk.
Can't deploy the fix, because
- developers trigger deployments through slack and I don't have access to the underlying deployment system
- infrastructure guys who have access aren't responding to my emails
Or a least document how to call the slack bot manually. (assuming it's just a http endpoint)
I'm good at not panicking about things I can't change, but I worry about some of my colleagues who find it difficult to not have control in these situations
I can't do anything to help them at the moment, so for now I'm heading to my couch with my analogue book :)
Do you honestly think a self managed solution or open source solution would be more reliable for most companies?
I’d still rather be on Slack and suffer a day of lost productivity than force people to use only email or IRC.
EDIT: And down goes Notion, too: https://news.ycombinator.com/item?id=25634159
What does this mean? What do cloud providers do when customers scale down their services? Do the providers literally power down servers? Do they sell the capacity to new customers?
I don't know if they power down some servers if usage stays low for a very long time.
See for example, Amazon Prime day:
https://www.cnbc.com/2018/07/19/amazon-internal-documents-wh...
Just a theory though.
Also, Slack has significantly more users in the US than in any other country[1], and it really isn't even close. So the offense you're taking is unwarranted anyway.
See page 12 of the document (which is page 14 of the PDF) https://d18rn0p25nwr6d.cloudfront.net/CIK-0001764925/70df834...
- Revenue is not same as users. Slack have tons of free users and some countries also has lower priced plans.
- Many companies like Amazon etc. probably is counted as US revenue for Slack but they have more than 30% of their employees outside the US. This should not be huge numbers but significant.
https://www.youtube.com/watch?v=UaPkSU8DNfY
GILORA: Starfleet code requires a second backup?
O'BRIEN: In case the first backup fails.
GILORA: What are the chances that both a primary system and its backup would fail at the same time?
O'BRIEN: It's very unlikely, but in a crunch I wouldn't like to be caught without a second backup.The Forsaken (season 1 episode 17)
LOJAL: I've been reading the reports of your Chief of Operations, Doctor. They gave me the impression that he was a competent engineer.
BASHIR: Chief O'Brien? One of the best in Starfleet.
LOJAL: Then why aren't the backup systems functioning?
BASHIR: Well, you know, out here on the edge of the frontier, it's one adventure after another. Why don't I escort you back to your quarters where I'm sure we can all wait this out.
Rivals (season 2 episode 11) KIRA: My terminal just self-destructed.
DAX: What?
KIRA: I lost an evaluation report I've been working on for weeks.
DAX: Even the backups?
KIRA: Even the backups.
There's a reason to have a backup to the backup by Destiny (season 3 episode 15)I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.
Every service is marked as "Outage" as of now (also when I wrote the comment).
If their service is down, you will have trouble. The service will be absolutely inaccessible. Don't give people hope with "may"
Another classic.
Even the major clouds...hn is going wild about it yet the dashboard says all good.
That apparently was also the case here. I started having smaller connectivity issues before it went down completely.
I assume this doesn't happen all that often.
1. (of a person) surprised and confused so much that they are unsure how to react. "he would be completely nonplussed and embarrassed at the idea"
2. INFORMAL•NORTH AMERICAN (of a person) not disconcerted; unperturbed.
I wonder if this is because I haven't used the phone app in a few days, so I was already logged out, but you and others were still logged in?
I imagine their infrastructure to send push notifications is decoupled from their infrastructure for chat services themselves.
It'd be interesting to know if they have a master switch to disable notifications in times like this where they aren't usable anwyay.
It's not like there aren't alternatives. You could even imagine someone has a live bridge between Mattermost and their Slack team, making the switchover seamless.
* lots of users are coming back to work after the holidays today
* lots of users take the holidays off and fully disconnect
* significant new users added in 2020, with so many teams going remote
Sounds like a possible recipe for infra scaling issues and/or cascading failures to me
EDIT: according to a random quora post they do, so keep the tinfoil out!
My response is if it’s important long term, it needs to be somewhere visible and exportable should the platform change. As it is now, Slack exports are horrible and large.
H&R Block's page there has this as the most recent comment, from an hour ago:
> My sister was able to have a bank pull money off her card yes her old card dunno which bank ill find out in bit
Two hours ago:
> I went to atm and thought I was crazy my pin wasnt working.
These reports are entirely useless.
Wonder if that's related?
I.e., it's down.
(And if you're saying that according to the legal blah blah blah of the SLA that this isn't technically "down", then there might as well not be an SLA.)
I am because Ive had these exact conversations with cloud hosted providers/products. Never once have we been refunded according to the SLA in our contracts. Never really down (according to legal).
I guess everyone hopping back online over the course of a few hours for the new year is too much to handle!
Potentially related?
So its probably a wider issue affecting everyone - network level is my guess.
It would be nice if they could fix it so that a fresh start also goes to that page, at the very least.
- Jan 4, 10:14 AM EST
The status for messaging and connection services has been marked as [incident]
> We're continuing to investigate connection issues for customers, and have upgraded the incident on our side to reflect an outage in service. All hands are on deck on our end to further investigate. We'll be back in a half hour to keep you posted.
- Jan 4, 11:20 AM EST
> There are no changes to report as of yet. We're still all hands on deck and continuing to dig in on our side. We'll continue to share updates every 30 minutes until the incident has been downgraded
- Jan 4, 11:52 AM EST
HN is also pretty slow...
Sending big hugs to their ops team.
For their app to just go completely offline is unacceptable. Bugs and degraded services I get. But this is catastrophic.
Doubtful it's a code issue causing a total system outage. I'm assuming they have a bunch of auto scaling infrastructure that wound down over the holidays and couldn't take the spike this morning.
If a thousand small companies have thousand customers each. And these small companies experience an outage per quater, then a million businesses suffer every quater.
As the end-user-business, is it better to suffer the outage at the same time as other businesses? Is it worse?
Surely there are valid arguments against relying on big companies, but I don't think this is one of them.
Not all companies are created the same. Microsoft, Google and Facebook have had their outages, but IME much fewer than Slack.
If there are a thousand small companies, none of them have a network effect, and those that experience more outages per quarter will lose customers to those that have less outages per quarter. So they have much more incentive to improve.
Whereas network-effect beneficiaries like Facebook (and to a lesser extent, Google, Microsoft and Slack) have much less of an incentive to improve. Who else would the customers go to?
Somehow Slack is very resilient in general. I also appreciate its UX/UI being far superior to Teams.
Ultimately, the cloud is often a single point of failure that companies become over-dependent. So I'd favour a free (as in freedom) and open source self-hosted/deployed alternative if there was one (even if it was from Slack and for pay). I agree with most on here that there isn't such a thing yet - but it's well worth building! So those of you out there who are considering implementing "yet another text editor", maybe this is something to work on.
I hope matrix/element will rise more.
* iMessage was taking its sweet time sending a few texts this morning.
* I had momentary trouble trying to call a business from my Verizon phone, and someone I know had trouble calling from AT&T.
Could just be a coincidence, but I wonder if something larger-scale is happening.
* Todoist MacOS app is having trouble talking to its API
Todoist was having issues and iOS app launching from Xcode started taking a lot of time in the middle of the day (which reminds me of the app online check fiasco not so long ago).
Even HN seems to be a bit slower.
Although it's built as a live news discussion site versus a team messaging app, the topics can be about anything, are public, and inviting others is as simple as sharing the url of the post (mobile/desktop web).
Example (reposted this hn post to sqwok): https://sqwok.im/p/Q3-1AZFLCSpjew
[0] - https://twitter.com/SlackHQ/status/1346132040249470979
I wonder if it's an AWS region issue
> We're continuing to investigate connection issues for customers, and have upgraded the incident on our side to reflect an outage in service. All hands are on deck on our end to further investigate. We'll be back in a half hour to keep you posted. > Jan 4, 5:20 PM GMT+1
It's also extremely bad that we're 1 hour in, and they are still "investigating", with no more details than that.
Slack goes down so often we're thinking of writing a very boring clone that uses ActiveMQ and MySQL, just because chat should be boring and needs to "just work".
For something so simple, you have to run a massive server, like gigs of ram and multiple core, even with a very modest user load. Take a look at the codebase, it's also a mess and impossible to fix any bugs. Finally, if you want to get your data out or report on the message activity, good luck, you'd be better off passing paper notes around. The open source version is nerfed a bit too, no LDAP authentication for instance, so it creates a lot of problems there too.
https://twitter.com/louiechristie/status/1346213038924427265...
I think HN is hiding these posts. Maybe status threads are discouraged now? But they're much more useful than status.slack.com etc.
They always have been, since they clearly don't fit the guidelines for what a good submission is and usually leave little for interesting discussions. (unlike postmortems of past outages, which often are good)
We're continuing to investigate connection issues for customers, and have upgraded the incident on our side to reflect an outage in service. All hands are on deck on our end to further investigate. We'll be back in a half hour to keep you posted.
Jan 4, 8:20 AM PST
Doesn't have to be, though. One person doesn't even have to tie their address to a single provider, and seeing past received messages doesn't even need internet connectivity.
Send help.
A plaintext web interface would keep my team moving along while they resolve their issues.
"Let's see, I'll look up so and so's name with Sla.... shoot"
"Okay, I'll just find that thing I .... nevermind"
Seems to be working intermittently, however.
I am really looking forward to a better competitor taking over their market share, I presume things will only get worse after Salesforce acquisition.
If you haven't been able to justify testing your PACE plan with your bosses lately, now's a great time to go ask again.
It actually might be a good thing that everyone doesn't feel the need to look at slack every X minutes.
And to be clear I don't mean Slack's implementation of threads which is hiding it away in a separate panel and which doesn't get used by everyone either.
Nice.
If a human created it, it can never have in human capabilities
edit: it is now showing as a total outage on the status page
Would be similar if auth was down. You can connect to us, you just can't authenticate so can't actually do anything.
Edit: Looks like they updated the status to properly show an across the board outage
/s
Edit: Not consistently, I guess. 9 out of 10 times it responds instantly, then it lags once in a while.
Depends on the timezone you're in, though one could theoretically cite a disparity between physical and mental/emotional/temporal time zones...