Slack – Degraded service affecting multiple features
status.slack.com
status.slack.com
Is there such a project, and if so does it have any traction in the real world?
(Oh and ditto for every damn thing Atlassian owns. Yikes.)
Have to say Mattermost's API docs are so much easier to read than Slack's.
Tells me there's some subsystem throwing the error, not the actual messaging (or, if intermittent, something wrong with the distribution.) What do you want to bet it's some kind of new analytics?
Typically a 500 should be considered as a "may or may not have succeeded", and steps taken to deduplicate operations if necessary.
/shrug
It has easy integrations with basically everything, and non-developers can set them up on their own.
It has bouncer functionality built in - You don't miss messages. It can also archive them, and handle them in a corporate-compliant way without needing to speed a bunch of time setting that up manually.
Because it's a central company and not IRC push-notifications are easy.
Emoji are shared corporate wide, anyone can add them.
It's a really nice out of the box experience.
You can absolutely replicate it with enough setup and tweaking, but the free version of Slack will go a long way.
It's not "better" than email, it's different. Real time. It has a million and one service integrations, which IRC doesn't.
Point is, it's really not all that crazy to just pay a third party to manage all this stuff for you. You might imagine the alternative is a perfect IRC server with 100% uptime, but I doubt that would be the reality. If you're saying a service like this is too important to be run by a third party then you're going to have to invest a lot internally to come up with an alternative that's more reliable.
It's IRC. On modern hardware. It's practically infinite.
> it has a million and one service integrations, which IRC doesn't.
There and were are things way before Slack, eg https://loqi.me/
First "but it doesn't scale". Now "but integrations".
There will always be another "but", and there will always be a link on the internet showing that it's been done already.
"but shiny web gui" - https://thelounge.chat/
"but mobile app" - https://github.com/MCMrARM/revolution-irc
"but it's complicated" - https://www.mirc.com/ since 1995
IRC does have limits. Some. One is that to make it "modern" it needs a bouncer - that said znc is spectacularly simple to set up with it's web interface.
Integrations were never one of those limits. Just google IRC bots.
I think it's really sad how the world has migrated away from an open protocol to a closed one, but I remember IRC having some very real annoyances, well beyond stuff like integrations.
That's my entire point. It's death by a thousand cuts. Now I'm responsible for an IRC server, a web frontend, a mobile app and a bouncer.
The OP said "just run an IRC server". There is no "Slack or just an IRC server". There's "Slack or a number of connected, interdependent services you need to maintain". I'm not saying it can't be done, or that it isn't the right answer for some people, but it's also a lot more work (and thus, expense) than just paying for Slack. It absolutely makes sense that people do it.
When did this madness start?
Paying experts in their field (chat tools in this case) is most of the time the most reliable and cheapest variant. You must also include that Slack has plenty of features that are not available in IRC and mean the productivity is not as good.
30 years ago, efnet was (apparently) 35k users.
Get 2 sysadmins. 2 because of redundancy. Pay them to maintain your service.
"well funded sales team and corporate support". F that. Trust your employees.
Yes. Someone above have mentioned ibm using slack though. I think that's huge enough for the argument.
Do you stop trusting them then?
If you want to match what Slack offers you aren't just running an IRC server, you're going to have a small fleet of dependent services that, added up, just about match what Slack does.
Is it really that surprising that people just opt to pay for someone to deal with all of this for them?
[1] - https://thelounge.chat/
[2] - https://thelounge.chat/docs/configuration#ldap-support
Now, modern IRCds might have features I’ve never experienced on raw IRC... but if they don’t, I don’t suppose you’re suggesting that this server also run thousands of bouncer bots in order to provide this contextual history feature?
(Also, whence push notifications?)
I can set you up a resillient architecture and 3 services (communication, documentation, filesharing) in less than two weeks (let's say two weeks and a half because security concerns/limitations and all that stuff). Add one week and you will have mail server, VOIP communication, log conservation and monitoring, backup images on CEPH, vnc service for "easy" maintainance and a small documentation explaining the few action needed to maintain the monster for two years. Initial cost for a company: 15 to 25k$, + 1 hour and ~100$ a month.
Truth is I see many capable people buying into these as well because of the convenience - read: lazyness - factor. The "ain' nobody got time for that" effects everyone and the willingness to pay for services had skyrocketed in the past decade. This applies to blogging, to home servers, backup, anything. We're all guilty at some level.
- no audio / video
- no history
- no SSO / AD integration
- no rich features like links, pictures, upload docs ect ...
- easy bot integration with APIs, I can have Datadog pushing things to Slack ( alerts, graph... )
- no archiving channels on the go, Slack is very useful when you have an incident and you create a temp channel and then archive it when the incident is closed.
There is of course several other server implementations in progress but none are fully ready yet.
The features you're listing is server/client dependent.
There are products that are far better for audio and video communication than Slack. At my company, we use Zoom.
> no history
Can be done with a bouncer if you need it.
> no rich features like links, pictures, upload docs ect
Is there a reason why an external service can't be used? I can't upload a image here on Hacker News, but I can always include a link to it.
> easy bot integration with APIs
So I have to use the HTTP protocol to manage a bot rather than checking for certain key words in the message itself.
> no archiving channels on the go, Slack is very useful when you have an incident and you create a temp channel and then archive it when the incident is closed.
That depends on the settings. One thing that happens every so often at work is that people get trapped as the last remaining person in a channel because only an administrator can archive a channel. In IRC, that's not an issue.
"The Wheel of Time turns, and Ages come and pass, leaving memories that become legend. Legend fades to myth, and even myth is long forgotten when the Age that gave it birth comes again."
* push-notifications that just work
* client available for enough people
* copy / pasting images that just work (the'll never admit it, but animated gif is their killer feature)
* stable-enough infrastructure for something that costs you 15€/month/user not taken from your "salary" budget (often a different and substantially less taxed part)
* and their outages are funny (receiving messages twice is arguably better than not receiving message)
Yes, it's because it "just kinda works for what it costs, and you can forget about it most of the time and do your job."
This seems to be characterized as "laziness" by some people, a point I honestly don't understand (I though lazyness was a virtue in our circles.)
That being said, there are legitimate reasons to not use it:
* software licence
* data property
* technical implementation of the native clients
* general distrust of centralized solutions
* cost of having IM at all (see Newport's whole publishing career)
At least, it means there's an opportunity to create a company that handles decentralized user-friendly open-source standard-based data-respectful attention-preserving uber-reliable chat system. HN users would be a market.
For some definitions of "just" and "working". Even ignoring the business limitation, the UX of Slack search is meh. That said, I'm yet to see any of the modern web&mobile-oriented IMs that would have a decent search UX.
But yeah, your list is solid. They offer plenty of value, just mixed with plenty of negatives.
A big source of the complaints against Slack is usually the network effect - in context of both work and OSS communities, it's usually somebody else that makes the decision and imposes it on you, forcing you to run face-first into all the "legitimate reasons to not use it" you listed. Half of the pain would be gone if they opened up their service to third-party clients; I'd say the reason they didn't have as much opposition from people initially is because they run the bait-and-switch with IRC gateway.
For work, the decision of which IM tool is most likely going to be a corporate one, without much choice of solution. I guess the problem here is the same as for other corporate tools. (And people have strong feelings about corporate tools, but it seems there are less strong feelings about "using jira vs bugzilla" as there are "using slack vs irc". Or maybe there are as much strong feelings.)
On the other end, I get the troubles with the "network" effect for out-of-work communities (open source projects, clubs, etc...) Such communities can at least have "some" level of debate about which solution to use, and balance the pros and cons as they see fit. It's completely acceptable for an OSS community to discourage using such platforms - as any other non-free, centralized platforms. We're still not forced to use anything.
Then, there is the "club of non-power user" situation. I would be hard pressed to ask someone from my drama club to setup and maintain an IRC server (neither do I want to do it.)
Thanks for adding this dimension.
Follow-up: is there a "nice" provider of some slack-like solution (mattermost, etc...) that hosts servers for communities of non-techies for not too much money ?
- Search
- Scrollback/history available in the main client interface people who newly join a channel (extraordinarily important during incident response)
- Ability to link to a message in history
- Ability to handle multi-line messages in a vaguely competent matter
- Image uploads, snippets, etc.
- Users of a GUI tool / web interface are first-class, users of a command-line client are second-class; this is unlike IRC, and important when the goal is a single chat system for the entire company, not just engineering (and usually a subset of engineering at that)
- Client/UI expectation that you do not catch up on every message unless you want to (like IRC, but unlike email)
- Quality integrations that have already been written with the third-party apps that you're already using
IRC and email have both been time-tested, and have failed the test.
I have seen enough corporate mail servers going down, irc and jabber servers breaking etc.
Using those won't magically give you 24/7 with more than five nines uptime.
Key difference is: If slack goes down many companies are affected and it goes on the internet. If a company's mail setup goes down employees are a bit annoyed, but after a while it works again and nobody outside notices and it's no news on HN (unless it's a real major thing bringing a notable company completely down)
Edit: clarifying that I'm talking about my workplace, not a company I run.
Unfortunately, Slack does not do this very well. I'd really like to know if there's another service that _does_ do this well.
For all people who think slack can replace email, think about these safeguards.
Also don't get me wrong, you do a great service. It's just a pet peeve that it seems invariably status pages are a lie.
”"”sauldcosta 4 days ago [-]
We use downdetector.com because status pages tend to take up to an hour or so to update, if they ever do. reply
jgrahamc 4 days ago [-]
1042 UTC First alert of global traffic problem 1057 UTC Internal group chat room up and running 1102 UTC Status page updated So, first alert to status page was 20 minutes.”””
In those 20m we had repeatedly checked your status page, realised it was our issue and started pulling engineers to deal with it as per procedure. People are on call, it's highly disruptive.
Surely you knew within those 20m that something was up?
In the end we realised it wasn't our issue because we checked Twitter.
Edit - added quotes
But, I guess we could have put some status up quicker.
I'm irritated, however, that you picked on Cloudflare as a bad actor here when we strive to be transparent and quick to get out full information whereas others (e.g. Amazon) are slow as mud.
But I get that being on the other side of this is difficult when you don't have information.
If it's any incentive, if the status page had any inkling that something was wrong, I'd be first on here posting "omg their status page is real".
That's the symptom of the outage as far as I can tell.
https://twitter.com/slackstatus/status/1144577107759996928?s...
It's odd because the message said it failed to edit, but my client still shows the message as edited.
If Slack is your only forum of communication with your team it’s time to rethink your support structure and DR plans.
Edit: It’s very unhelpful if I point out issues and don’t provide solutions. PagerDuty has a great article about Incident Response[1].
Ideal would be a live conference call (phone, Skype, Hangouts etc...), but you've got SMS, email, heck, I once solved a pressing issue with my team in a Sharepoint Word document where all our comments were just new paragraphs being edited live (which actually had the handy benefit of being savable without needing to pay extra for the privilege...).
However, when major backbones go down it takes it a while to recover.
[1]: https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...
We build services on top of that that don't cope with failures.
At the end of the day, that sounds like a massive pile of SPOFs that just barely keep up..
As app is 7 but we have grown beyond the seven into meta layers of service types which have become the favric of how we operate... such as email, txt, chat etc to keep all other aspects of infra up and running.
Also, everything above layer 4 is a enormous mess, considering that stuff is usually happening at the application itself anyways. There are network protocols which have proper OSI seperation, but they are rarely used. (IS-IS in combination with CLNS for instance).
A failure of mine was to not keep a project portfolio runbook which i could look back upon.
Imagine a system where all your status reports were logged and searchable - that would be move valuable than linkedin...
Sure, having a single connection is a SPOF, but anyone living in an urban area and anyone with a cellphone (and not using the same cellular carrier for home internet) has more than one connection.
A lot has to go wrong for me to not get work done. If Slack goes down, I use text, email, or Jira comments (in roughly that order of priority). Relying on only one communications method is foolhardy.
If there was a lower reliance on IM for all tasks, other solutions for crisis management would be utilised (potentially).
> failing to address pending scalability issues
How is that not related to not getting enough funding and resources?
And yes, having scalability issues means moving too fast - the _business_ moving too fast.
* do not fit well for a quick message interchange like you can have during an outage (email)
* don't have the whole company (or at least the whole tech team) on it (whatsapp, telegram, you name it)
* make compliated to share big chunks of text/logs (phone call)
It probably makes sense creating another secondary - usually dormant - communication hub like a WhatsApp group for emergencies but people tend to misuse it (like using it to ping for some minor incident when all the other more suited communication tools are working perfectly).
Anyway it wasn't a big outage on our systems, just a minor hiccup, but it's a lesson learnt nonetheless.
You won't have access to the things slack integrates with but at least you can talk.
I know some people use IRC as a low-dependency fallback channel.
I had an ircd of some description installed, configured and ready for use in under 3 minutes, and started to get my team using it as a stop gap measure. Had that rolling long before either the outage was resolved, or the HipChat server was resurrected. People forget just how stupidly easy it is to set up IRC.
Freenode runs this one for a server.
[1] - https://www.unrealircd.org/
[2] - https://www.anope.org/
[3] - https://thelounge.chat/
I had a similar outage at another job for a day. Pre slack days, but instant messaging and email was down.
I needed to do some things so I just stopped checking with people. So did other people.
The result was I found was that we were double checking, coordinating and doing a lot of verification that honestly wasn't needed. The sky didn't fall, everyone still saw the changes and were ok with it.
After that folks stopped doing a lot of the traditional coordination that proved to be superfluous and maybe never accomplished anything. The handful of mistakes that happened, were also easily caught / fixed.
Granted, this requires people to make good choices.
More study are needed to clarify if this work-related communication might actually be required for work to get done. (Cal Newport is definitely preparing a book on that any time soon.)
Unfortunately, the researchers are too busy maintaining their IRC server to actual do it - but at least they're using IRC.
Whereas, my productivity definitely goes up when HN is down ;)
* update the version of the IRC server
* update the version of the os of the machine running the IRC server
* repair a broken disk / fan / overflowed disk on the machine
* add / remove / reset password / change weird settings of users
I'm not saying this is "unbearable", and plenty of organisations have people whose job description would probably correspond to doing those tasks.
But I can definitely understand why you would want to skip them entirely and have someone host your chat - which is basically the job slack is paid for.
"Mission critical, but not paid by your customer" is always tricky to staff for, isn't it ?
Odds are that you’ll have updates worth installing once every couple of years.
>* update the version of the os of the machine running the IRC server
Almost never unless there’s a remotely exploitable code execution vulnerability. Local bugs wont matter unless you run multiple services on the same box.
>* repair a broken disk / fan / overflowed disk on the machine
Depends on your hosting setup. With a cloud setup perhaps never.
I’m certainly not trying to suggest that anyone should use IRC over slack in any situation, just that it’s not a horrible time sink requiring significant maintenance.
In this forum you'll be hard pressed to find someone who thinks it's a good idea to leave a cloud-based (or any internet-facing, really) server completely without maintenance for any significant period of time.
If you had said that usually it's as simple as checking in every now and again to reboot / make sure security updates are installed / check functionality and that it's a manageable burden which is worth the effort you might have gotten a more favourable response.
Edit: or if you had automated it to an acceptable level, maybe some more information about that.
Also, how many IT service desk ticket systems just lit up with people asking them to fix Slack?
Best wishes to all of the folks at Slack and those who depend on it.
https://status.slack.com/ibm/2018-03/f01d4c22cd953dd7
https://i.imgur.com/Rk6Kdgp.png
EDIT:
also that channel made using slack impossible for mac book air users, they had around 80% cpu usage for slack. so basically entire marketing and PM part was unable to work that day. developes machines were wasting around 20% on slack.
after people started complaining in that channel, posting in it was limited to admins only, but they didn't lock commenting. so, for approx 6 hours all of IBM was posting memes in THETHREAD as we dubbed it, and @mentioning the genius who created that channel. next day the channel was nuked, not even an archive preserved.
some guy calculated that the entire affair, considering electricity prices, a 20% decrease in developers productivity, and so on resulted in IBM loosing several millions with that stunt
fun times
p.s. shout out to Martinj for that spicy jeff-coffee-mug meme
Yes, you can mute a channel and suppress @everyone and @here if you want. Still friggin nuts that a chat service can bring a modern machine to it's knees.
Learn something new every day.
There was one specific person who would ask a question, expect an immediate response, and if they didn't get the response they would @here. They were instructed not to do that because it pings over 1000 people, so the next time they asked twice in an hour and when they didn't get a response they said "I hate to do this but I really need a response @here". They were kicked from the channel. Really unfortunate since it's a very useful channel to be in, but @here and @channel is just so disruptive in big channels.
Are you saying it is the user's fault where it is clearly Slack's inadequate implementation that is the problem?
Software designed for large numbers of people shouldn't go down because one of those people does something stupid.
The real fail here isn't putting everyone in one channel. The real fail is not using/providing proper admin tools
In general I find Slack not very concerned with your own ability to control your experience (no ignore user feature, can be added to rooms, can be added to rooms that then get pinged with @channel and your notifications start blowing up, etc).
This may be 'expected' in a small company where of course everyone should talk to everyone, but for larger companies and the plethora of various open-source or interest-based communities now using Slack, it's not so good.
Manually having to follow threads and unfollow threads is a nightmare compared to the one time configure experience of channels, especially since you cannot just mute all threads, and is easily the worst part of Slack to me.
> 200'000
Interesting. Why an apostrophe?
1600000000
1'600'000'000
separating at 10³ is due to counting customs of the region i live in.using an apostrophe `U+0027` appears to be the least wrong use of possible glyphs accessible on my keyboard layout
also c++14 does it this way
I understand that post mortem's are interesting learning lessons, but if you're curious whether S3 or whatever is down or not right at this moment, then use the subsequent status pages of those services.
It's also important - many here use Slack, and an outage may be confusing - most of us don't immediately look at status pages, so we have somewhere to go and discuss the happening.
Besides, the discussion here is useful. Someone has already pointed out that some messages are getting through despite returning 500s, so integrations are spamming duplicate messages everywhere. That's useful to me if I'm running an integration.
(besides x2, it's one thread with a very self explanatory title. Maybe just don't read it?)
If you don’t think the article is appropriate flag and move on.
This community is awesome. Always someone with their finger on the pulse :)