Users may be unable to connect or experiencing degraded performance
status.slack.com
status.slack.com
* Messages are shown as sent (not grey), then, after several minutes, a "Try to send again" message appears. Or the message turns grey. Or both.
* The webapp randomly reloads. Sometimes already typed messages are preserved, sometimes they disappear. ( I assume a reload happens when a new version is deployed with some "force reload" flag )
* Threads and channels just show a loading spinner forever.
* After switching channels, all fonts disappear and just the menu background remains. No remedy except reload.
Outages happen. They get resolved. That's live, and it's fine.
What is not fine is how ill-equipped the frontend is to deal with such failures.
I have been seeing some version of this on my laptop for months. Don't seem to have the same issue on my work computer. Maybe it will finally get resolved.
But, how do out count downtime of a distributed system? We probably need a new way of discussing availability.
An outage to part of your system might only impact one or a group of users, but to those users you could have been 100% down, effectively, even as other users were not impacted at all.
So maybe a series of user/%age counts? "This quarter 90% of our users/use cases had 100% availability", for example. With the gold standard being 100/100. Maybe numbers for 99, 90, and 50% of users/use cases? IOW: "1% of our users had 95% availability, 10% of users had 99.9, and 50% of users had 100%."
# units successful / total # of units
The converse is pessimistic rendering: waiting for the server response before updating the UI (and possibly showing a loading spinner in the interim).
Facebook Messenger has the best UX paradigm for this I've seen, since it's most graceful in handling the edge case failures while still presenting as optimistic. Your message immediately appears as a conversation bubble, with an icon displaying if the server has received it, an 'X' if it fails (clicking the icon gives you options: re-send, delete, etc), and a read receipt for when the other party views the message. It's a fantastic way of handling the lifecycle.
Slack fails because it doesn't communicate that pending, the-server-hasn't-received-it-yet status, and the fallback isn't as graceful.
I hate the phrase "I'm a realist not a pessimist" as much as the next person, but isn't this actually "realistic rendering"?
"Optimistic" often means "assume X, perform or calculate something, change the result if X turns out to be false".
"Pessimistic" often means "assume not-X, perform or calculate something, maybe change the result if X turns out to be true".
This comes up in other domains.
For example in compilers (and reverse engineering decompilers ;-) there are optimistic and pessimistic analyses.
An optimistic analysis might initially assume a particular variable has a specific value at a particular location, or a branch is always taken. Then trace through the program paths, using that assumption (which limits those paths). If it loops back to the location with a different value in the variable or different value for the branch condition, broaden the assumption and update the trace.
A pessimistic analysis might initially assume the variable at a particular location could have any value because it hasn't yet traced everywhere to find out the possibilities, or initially assume both directions at the branch could happen and include them both in the trace. By the time it finishes tracing all paths and values, those assumptions may turn out pessimistic and it can start pruning away unused options.
These two strategies can yield different answers, and for some analyses the optimistic strategy gives more accurate results than the pessimistic strategy.
So which one is more "real"? :-)
FYI, Signal has used this paradigm for quite some time (I'm not sure which software predates the other on this feature), it's two checkmarks which remind me of the old "open Apple, closed Apple" keyboard keys - message bubble is shown with two open checkmarks right away, a single closed checkmark means the server has received, the second closed checkmark means the other end has received (assuming read receipts are enabled, Signal allows them to be turned off for privacy reasons).
You post a message, it looks good, you move on.
Then later on you come back to it and it was never sent.
The only way I found to know that your message has been received is to have an ack from your interlocutor.
Viber had a similar thing long before FB's messenger: single gray check mark character if the message is sent, double gray check mark if it is received by the recipient and double check mark in purple if it is seen.
The only issue is with the purple color of the "seen" as it can sometimes blend into an image behind it (if the message was consisting of an image).
It's easier than making the UI and server faster for sure, I'm just not sure why engineers don't insist of fixing things the right way.
Einstein figured out why in 1907, and most of us have just been going along with the assumption that superluminal communication just isn't possible.
In majority of cases, I assume network latency will be the dominant factor, and that is outside the control of the engineers.
I happen to use slack from hardwired gigabit ethernet to a gigabit fiber internet connection. But I maintain some awareness that some others may use it from their older android phone with 2 bars of service, or on a laptop with 27 open blasting wifi spots, etc etc etc.
I will therefore put forward that assumption you can ensure instant communication, is the wrong wrong wrong way for engineers to write code...
That's like saying "just fix global warming", it's more complicated than that.
There are other things to consider too. Like for example network failures vs backend failures. In the case of network failures you can actually notify users before they even try to take an action. So now you’re down to some diminishing case where your code is effectively wrong (maybe the client is too old), and again there are more elegant ways of handling that. Also, I bet you have the response back really quickly to make a call on it.
Having a worked on unstable networks I can say that there are parts of slack’s implementation that suck - like if you try to change a message that hasn’t synced yet, they get into broken states with the edit. But that’s not the fault of optimistic UI, it’s other technical choices at play.
I know you haven’t named a horse in this race, but I feel people hating on Slack for having an optimistic UI are a bit off base. Considering that I’ve sent (apparently) 2.5k slack messages in the last 30 days and I don’t recall any going wrong, I think they’ve got it right.
There are trade offs to consider in UX. You’re trying to make things as seamless as possible for the end user without distracting them with things they don’t care about.
Overall I strongly dislike this UX pattern. Users should have immediate feedback about actions that haven't completed, but that feedback should only tell them the truth: that the operation has begun.
It's not a 1-3% failure rate for the backend, it's a 0.01-0.1% failure rate for each request between dozens of microservices each with their own bespoke error handling, each owned by a team with more political clout than the frontend guys/gals who keep reporting Sentry errors due to "UnknownNetworkError" or some other such nonsense, all wrapped in some spaghetti code frontend thats written like the whole thing is running locally.
I somewhat often send messages just before I have to run to catch the subway — and when I think the message got sent, I type `sudo poweroff`
If I post a message in a channel and go back to work, I don't want to find out later when I check for replies that my message didn't actually send - I need a notification!
Even if the idea is to fool the user in to thinking Slack is faster than it is, the UI could at least use the reversed 'light grey until confirmed' UI once it's detected that there's a problem. That would make it a bit less tiresome.
At which point there is no option to cancel and say "you know what... don't send it".
If you notice it 24 hours later, your choices are (1) be nagged about retrying it forever or (2) retry sending a message that may now be irrelevant.
When it comes to retries, the easiest thing for the user is for the software to be responsible for retrying failed actions. Choosing not to do that makes sense IF the goal is to give the user more control because they may not necessarily still want the action done.
But if you don't give them that control, it's the worst of both worlds. The retry happens slower, with more manual work, and without the benefit of choice.
On mobile it means that I can't tell if this user has responded to a message or is online at all. I tried telling it not to send ages ago when it actually failed but it hasn't worked, and frankly now it's so far back in history it's impossible for me to try again.
I've deleted slack and re-installed to no avail. It's really an awful problem.
(I know this because I just spent 60 seconnds staring at my message trying to figure out why the formatting was off before it told me it wasn't sent. It sent immediately when I retried though.)
This morning alone, I've had multiple messages start off black, then turn grey after about 10 seconds. And that's not just true today, but other time Slack has been down as well.
"If it's black, it's confirmed delivered" is not an invariant.
Recently i've realized that this actually introduces more typos -- where a large number of phrases are duplicated or corrected typos reapplied.
There's something wrong with slacks operational transform or CRDT implementation or whatever they use to get edited messages to converge
I'm pretty sure they have none of that, considering you can only edit your own messages. So what's likely happening is that two quick edits are reaching the server at the same time and overriding each other.
I suppose it's possible that the corruption happens on the client side tho --
1. initiate edit action and get editable view of snapshot of text
2. server acks some previous edit request causing some client state related to the message to update and promote a previous edit in some way
3. hit save after finishing editing -- and some delta operation on the client side updates a different message state than the one i initiated edit action from ...
But that doesn't mean my observation of the behavior of the bug is unlikely.
And also -- there is not a hard requirement that only one editor exists -- the same user could use multiple devices at the same time. Consider that and you might find your assumptions to be invalid.
(Where a message sent doesn't match what was typed sometimes.)
The RGBA value for the "unsent" text is #1d1c1db3. Normal "sent" (or "attempting to send but haven't yet decided sending will fail" is #1d1c1d. The RGBA value (alpha = 0.7) ends up being equivalent to #616161 when rendered over white background which the Slack default skin does.
Demo for reference: https://jsfiddle.net/9smqe8y7/
I don't mean to disparage at all, but IPS monitors are far from universal.
Just... how!?
Basically time "down" / total time elapsed during $period. At the beginning of the month/quarter/period you'd be at 100% (if they are up at that moment). Or even simpler, just have it be a measurement of a rolling window X days in the past.
I've questioned status pages since.
Edit: Not just for slack, but 2020 in general
Edit: Folks, I was kind of half joking and making fun of NDT. And we're discussing the definition of "Arbitary" and harvesting seasons. Lighten up people.
Did Neil DeGrasse Tyson really say this? Seems very close-minded, as if physics is somehow more meaningful than the only thing that actually makes it worth anything at all.
Life.
"arbitrary" defined as "based on random choice or personal whim, rather than any reason or system". Since a year is one revolution around the sun, our source of energy, it doesn't seem arbitrary at all.
Between the rotations of the earth (day/night), revolutions around the sun, changing of the seasons, phases of the moon, we basically live in a clock.
I understand there are benefits and reasons for the state of affairs...
Between that and using vscode remote SSH onto my server environment, it’s almost like the old days of save-refresh, but with auto refresh!
We have build times in the dozens of minutes thanks to a decision in the past to use Dart.
Hey, so in 2020 we have unlimited sandboxed environments that do not require virtualization and can be spun up, destroyed very quickly, and their construction is fully automated. And it's somehow a bad thing?
Get off my lawn! When Docker arrived it was a godsend, given that we were trying to do similar things directly with LXC.
(I'm fully expecting someone to chime in saying, 'you had LXC? All we had was ones and zeros, and we had ran out of zeros')
Some interesting logs most people can understand:
[REDACTED] info: [OPPORTUNISTIC_RELOADS] User has been blurred for 1200000 ms, evaluating whether or not we should reload opportunistically
[REDACTED] info: [OPPORTUNISTIC_RELOADS] Client is still visible to the user even though it is not focused, will avoid updatingEDIT: problems == sustained issues. Yes, there are occasional blips.
The screen sharing one is so bad I have to send zoom links out if I think I’m going to screen share in a call
Today’s disruption should provide an interesting data point all of its own.
Regarding long term investment. For most people at most companies long term (more than 2-3 years) investment means moving to a better job somewhere else. Investing in your future is figuring out how to get the next job, not maximizing one aspect of your current job that your employer doesn't even seem concerned about.
I'm not delighted by their use of might/may language and lack of detail in the status messages. Slack is pretty clearly not working well for a large number of users, and based on the above IDK what can be attributed to degradation vs. people being unresponsive.
Also check their status page in a month or so, and you'll see "October 5, 2020" with some message like "no issues".
Assuming that a random company will pay to have engineering/ops staff on-call 24x7 who all have enough knowledge about their internal chat system.
Or pay a company to host your chat system, but use an open source server and just have them manage your instance. If something goes wrong, you email them and they can quickly fix the issue. Because it's a simple system and they know the problem exists on your node.
Besides, I don't need it 24/7. There's always email for that. I need chat during working hours. Slack is currently not fulfilling that need, and even if I could get ahold of someone there, that person wouldn't be able to fix my problem. Their system is too complex.
that includes deployments, updates, security audits, etc.
> By uploading, distributing, transmitting or otherwise using Your Content with the Service, you grant to us a perpetual, nonexclusive, transferable, royalty-free, sublicensable, and worldwide license to use, host, reproduce, modify, adapt, publish, translate, create derivative works from, distribute, perform, and display Your Content in connection with operating and providing the Service.
I'm not a lawyer, but that seems like a key limiting clause in that statement.
Those absolutely seem like overreach in terms of what someone would consider reasonable from a user perspective.
I assume Discord's gamer branding (which they only dropped recently) might be the reason Slack is more widely used for teams.
So let's say I wake up in the AM and there's a thread that says "how do I use the API for companyservice.example-api.com" and the thread has 50 comments.
Let's say it's a service I support. First, I know there's an issue because a customer required attention to begin with. This is then confirmed by the fact that the thread has 50 comments in it so there are probably documentation and usability issues with this service.
The company I work for and the services I am responsible for deals with internal customers and we use these two absurdly simple metrics to optimize support regularly.
When we first started with certain services that my team is responsible for, we used Slack metrics A LOT. As in, we had spreadsheets with this data and we'd go over them every week. Nowadays, our services are almost 100% self-service and we rarely, if ever, get pinged on Slack for those services because the documentation is optimized and the usability has been specifically designed so that our engineers don't get in the way of our customer's needs.
It's useful when people arrive in the discussion at different times and you don't want to necessarily let the whole channel to be notified about your specific remark, and don't want to write a private message either -- one reason being it would lose the context.
In our project we have multiple channels that have actively 3-5 persons and work at different times. The things we chat about are lightweight changes or news, or maybe someone has a quick clarifying question about implementation detail. Something that wouldn't be posted in an issue for one reason or other.
Remember the irc support channels where conversations and problem/solution dialogs were constantly muxed with other conversations / dialogs?
Imagine each question being a root message, and all of the dialog / suggestions / answers being in a thread. Also makes it super convenient for search.
In addition to that I find that beyond a certain channel / DM count it's a lot easier to follow up on things with threads
Give it a shot, try to get your whole team on board. You can always revert back to not using them later :)
I'll recognize that I'm a bit of an oddball in that respect nowadays, but it can work, at least with a small enough room. I love having all the conversations cross over one another and just being able to read back and follow them.
Whatever happened to the distributed in distributed systems?
Just how difficult is it for text typed up on my computer to appear on your computer.
We are facing this error for the last few weeks, in case if a member drops off the call due to internet issues he/she can't join back until ~15 to 20 mins. We all are pissed off because of this horrible issue and planning to use something else for our stand-up call
This is more like how their infrastructure currently looks like I guess as it’s somewhat newer and follows up as an update to your video.
They still use MySQL though.
https://www.youtube.com/results?search_query=aws+rick
is pretty good on why some NoSQL approaches are a step forward (perhaps not MongoDB at scale if consistency is necessary https://jepsen.io/analyses). In particular though:
https://youtu.be/hwnNbLXN4vA?t=992
There could be other issues about why Slack is slow. But at Slack scale, you need to be extremely heightened in your database strategy or you should follow the industry and use Cassandra/DynamoDB's built in partition tolerance. Key value stores scale horizontally much easier. B-trees don't scale as easily horizontally past a certain point.
Essentially, good NoSQL DBs have abstracted scale for you (so you don't have to think about it as much). But you have to know the access patterns in advance (the types of queries and updates you'll be running for most use cases), since you need to design your table around these access patterns. RDBMS leaks scaling from the abstraction (you need to use message queues, etc.).
You keep using MySQL at Github and Slack if you want periodic downtime/degradation in service tho.
I think you're mistaken about Twitter:
"Manhattan(the backend for Tweets, Direct Messages, Twitter accounts, and more)"
https://blog.twitter.com/engineering/en_us/topics/infrastruc...
I'm not sure about the others (e.g., there's nothing recent for YouTube I could find.. likely they'd use Spanner though for things like comments?).
I think you need a really talented database infra team if you're trying to use RDBMS for something like a real-time messaging store at scale. I can't say more about Amazon. But I just don't think it makes sense for these use cases to use RDBMS where real-time messaging (pull requests, comments, etc. for Github - messages, posts for Slack) is the 90% use case.
UPDATE. Maybe I'm wrong can you can use: https://vitess.io/ for horizontally scaling MySQL. I don't know enough about the details of it. But getting the data store right is so important to the overall backend's stability (I think it's no coincidence that Twitter stability became "solved" when they moved to something like Manhattan). And I just don't see why you wouldn't rewrite things using something that logically makes a lot more sense instead of trying to push connection pooling, query rewriting, etc. to the limit. They don't fundamentally solve what consistent hashing solves.
EDIT. Also, wrong about LinkedIn re: messaging:
https://engineering.linkedin.com/blog/2020/bootstrapping-our...
https://en.wikipedia.org/wiki/Voldemort_(distributed_data_st...
https://en.wikipedia.org/wiki/Consistent_hashing
Discord understands to use this too.[1] Not to say you can't have your user database in SQL like Facebook, etc. But for messaging? And really for anything with high throughput / low latency, where you know the access patterns, just doesn't make sense to not use something with consistent hashing.
But as mentioned in Rick's AWS videos about https://en.wikipedia.org/wiki/Technology_adoption_life_cycle - there's a lot of late majority, laggards, etc.
My original point - why I raised this to begin with - is I briefly browsed that Slack CTO's video. At one point he mentioned using RDBMS "because that's what we're experienced with." That's never a logical reason. It may be a practical one. But with time... it just doesn't stand up to ideas that are better and have proven themselves (e.g., consistent hashing). But again, using MongoDB the wrong way or assuming the "document store" is the main reason for using NoSQL can confuse people (it's one nice benefit for ad-hoc data models! but the big innovation in NoSQL is consistent hashing for the low-latency / high throughput use cases). And SQL has its benefits for certain use cases. But there's a better solve for the messaging storage at scale. I post this because: (a) I'm interested in others' opinions and feedback about how they've made RDBMS work (thanks) (b) tired of Github and Slack being down periodically, and for each mature SaaS to go through this learning curve (like with Twitter). Yo just use DynamoDB or Cassandra and save yourself the time/effort.
[1] https://blog.discord.com/how-discord-stores-billions-of-mess...
The slack CTO's comment about choosing RDBMS because 'familiarity' is interesting. IMO it's a gamble. I've seen it happen with my company when being a latecomer to containerization.
When it came to picking a container management tool, it was a tossup between k8s, Nomad, or just saying to hell with and running those containers ourselves on EC2 instances. Having run our stack on bare metal for year = we were really pretty good at it. There was a surprising amount of automation that could be ported over.
Eventually we picked k8s, and coincidentally, our usage grew more in 6 months than it had in the last 2ish years. So all in all, the gamble paid off.
... but I like to think there's another world where we picked the 'its familiar option' and things still worked out. If our traffic hadn't grown the way it did, we would never have felt the pain of having to manually scale out our systems - or basically write an in-house version of Kubernetes.
So in that sense, I'd guess that maybe some teams have the bad habit of playing the same side of the coin everytime. It may be prudent to stay conservative when picking a Datastore, maybe it's would be smart to pick a risky technology for your app servers? (and vice-versa)
No new updates for the moment as we're continuing to investigate. Customers may notice that other functionality, such as Search, is also affected.
Oct 5, 12:40 PM EDT
(hint: it's not great, but at least we can fallback to Chime)
Actual rollout started couple months later. Currently living in a nebulous/purgatory in-between state, most people have both Chime and Slack installed on their laptops/phones. Can't fully commit to just one or the other, yet.
Be honest to your customers, or I will make sure our team switches to Microsoft Teams.
A Teams "channel" has counter intuitive behavior, the messages you send are more like forum threads, users can reply to them, and when that happens it reorders the whole messages to put that one at the bottom. It makes it impossible to "catch up" on a channel since people will happily reply to any messages so the ordering is completely random.
There is also a completely counter intuitive distinction between "groups chats" and "channels".
Teams is also the only chat service I know where there is an actual limit on the number of non-public channels you can have. After 30 channels have been created, you're doomed.
Want to remove an existing channel to make room for a new one? Nope, you cannot remove a channel, only "archive" them. Archived channels are removed after 30 days and there is no way to force their removal early. So basically if you reach 30 channels you have to wait minimum 30 days to create a new one.
Did I mention the UI is slow as hell, and half the screen space is used by unnecessary large padding between messages? I am rarely able to have more than 3/4 messages displayed on a 23" screen.
The only positive point of Teams is integration with outlook 365 and meetings. It works really well, quality is good, reliable, nice UI, présentation mode, screen share, etc
Everywhere I've worked Slack's basically a free for all, can add custom emojis and create private channels etc.
Teams was always more locked down. One place even disabled the GIF feature!
I know you can lock down Slack too but nowhere I've been has bothered to do so.
I have had a few people on calls with poor connection speeds. It degrades horribly by stuttering and distorting audio. It performs worse than a phone call. The best fix would be for Australia to fix its woeful internet.
When you load the web interface it identifies each person with their initials, not full name. This isn’t helpful when you have never met the person before.
I've used teams in the past and it's not that bad. It's a valid alternative to slack imho.
I know I sound like a Zulip shill, but try that for a bit and you'll see the difference, I love it.
Recalling the incidents in September, it's a bit cheeky they call that "99.93%" uptime (30 mins of downtime).
Having remembered that there were at _least_ 2 seperate 1 hour+ incidents where I was struggling to send messages in Sept, I went through to look at the history, and they seem to have purposefully hidden these incidents on the calendar. There's no way of listing incidents in a month.
The information and accessibility of the status page is obviously politically motivated by marketing.
It's like when trains report 95% on time, except that the times when the service is delayed is when the trains are full. All the delays that actual people experience (person-hours) happen in that 5%.
For me it has only been slow performance all morning. Its not good, but not what I would call "down" either. At least from my perspective. I could have a bunch of other folks who are completely unable to send messages and I just can't tell.
>Some users may be unable to connect to Slack, while others are still experiencing general performance issues.
Slack is a not-great implementation of a good idea, including a use case that IRC and XMPP don't meet. Tulip is a better implementation of that good idea, plus yet another good idea.
Example: We pay for looker. One day I thought to myself I'll spin up Metabase and try it out for myself against our prod database. It worked very well, fast and easy to query my data unlike Looker.
I haven't used it in about a month and today I try to access it but the app just doesn't load. Now I need to spend time to look into why this app isn't loading. Nevermind that I have zero Java experience or devops experience. I just ran a Docker file per their instructions.
It's often cheaper to pay for a service.