Real-time messaging
slack.engineering
slack.engineering
One thing interested me: Why the difference in pathing between events and messages? I think the event flow makes sense, but why not have messages also go through your gateway server instead of through webapp? Surely there is needless latency there when you already have an active websocket open to gateway? I thought perhaps it was because your gateway was for egress-only, but then the events section made it clear you also handle ingress there.
AOL chatrooms in 1995 too were blazingly real-time even on dialup. Anyone ever use progz/scrollers? With an OH (overhead) account you could flood rooms with thousands of messages a second. Rich text, not plaintext. And it'd appear in super fast "real time" to everyone in the chat.
350K – https://www.theverge.com/2020/6/4/21280829/slack-amazon-aws-...
The Kubernetes Slack has close to 200K users right now.
Amazon has 1.5 million employees, and I'm sure the majority are on Slack.
Forget 1995, no service other than Slack can run group chats at this scale even today.
80% of Amazon’s employees are warehouse and drivers.
But really is the internet actually doing anything new? My grandfather used to send telegrams overseas all the time. Does anyone remember teletype? How is that any different than us chatting with each other here on HN? Phaw, tech bros just keep reinventing the wheel.
AOL existed outside of America.
my fondest memories as a would-be computer hacker in the early-noughts was sitting in a BT phone box in the UK at 3am using the lineman opcode for free local calls to connect to my nearest AOL pop, so I could use one of their trial CDs with a fabricated or stolen credit card (I deactivated the trial before the first payment, I was not a total thief).
regarding doing what slack can do, I think we do have a bit of collective amnesia about what was truly good and truly bad about what came before.
Slack does a lot of things, but its UX is actually pretty piss poor due to the bolt-ons (threads is just about the worst implementation I have ever seen), walled garden, weird account things (magic links~ if you want to see your other workspaces make sure you use the same email everywhere!), security (goodbye decent realtime bot api) and the client which feels slow despite using more resources than it should given its status as a program that needs to always be running.
To take one example of a worse system: IRC scales to tens of thousands of users on a single machine; doesn't have emoji or picture sharing though- but the bot API is simpler for sure, you can connect in one click if the person has an IRC client (or you pass a webirc link) - no signup needed.
I think the problem with slack is exacerbated in europe (or rest of the world in general), where the latency of fetching thousands of tiny assets or chats is many-many times higher than in California.
Also; the reason I had to do that was because my mum wouldn't permit a home internet connection; broadband was a thing at the time. Dial-up was not mandatory in 2002. IRC with dial-up was faster than slack is today too; you can argue that slack “does more” though.
I claimed multiple things. Unless you think the single assumption being wrong in magnitude (I still maintain the majority of their customers were American) completely upends the entire thesis, then you're trying to treat this as some counterexample which... doesn't work here.
> Slack does a lot of things, but its UX is actually pretty piss poor due to the bolt-ons (threads is just about the worst implementation I have ever seen), walled garden, weird account things (magic links~ if you want to see your other workspaces make sure you use the same email everywhere!), security (goodbye decent realtime bot api) and the client which feels slow despite using more resources than it should given its status as a program that needs to always be running.
I'm not discussing the UX here I'm trying to talk about what Slack is doing differently than AIM.
Once you were online with AOL, it was extremely high speed. Virtually instantaneous chatroom messaging with hundreds of users in a given chatroom.
Mobile devices matter because they cannot maintain persistent connections. This by design necessitates a different architecture. AIM and IRC can have a single machine hold open a TCP connection and send and receive messages. A mobile device's modem wakes up briefly, performs some network chatter, then goes back to sleep. It's a fundamentally different requirement on the way the system is designed. IRC has evolved (through add-ons, not in the protocol itself) to shim this by using bouncers or having some sort of push notification based proxy for messages.
The other question is, how reliable was AOL? I have no idea how many outages they had, how long those outages lasted, nor what the impact was. Given the service was meant to be used as entertainment I suspect their SLA was one of best effort rather than the measured kind that Slack is leaning into (wrt using multiple CSes and using consistent hashing to pick which CS to send a message to, so architecture can be scaled as needed.)
I don't work for Slack but I work on high scale systems for a big company for a living. I'm pretty aware of the limitations of networks nowadays and the differences between systems in the dial-up era from now. Just the fact that Slack accommodates devices with push based semantics (mobile devices) and pull based semantics (anything that can open a persistent connection) is itself a large difference between systems of old and systems today.
Btw the TOC and Oscar protocols that AOL used have been reverse engineered and documented online so you can see for yourself at least at a protocol level what's different.
This isn't really a property of the device. It's an OS limitation. Nokia Series 60 didn't have a platform push service, so WhatsApp had to maintain a connection from the app to the servers to get messages without delay; it was not unusual to find Series 60 phones connected for 30+ days when I worked at WhatsApp. The newest versions of Nokia Series 40 had push, but older versions had to long-connect as well. Some mobile networks aren't great at letting connections work for longer times either, but some can.
But yeah, Apple only lets Apple keep connections open. Android doesn't have hard and fast rules, but background connections are likely to die, you should use push if you can. The push services connection is going to try to stay connected for a long time though.
Thought it is quite interesting to consider a world where mobile platforms made different platforms.
If you’re using the word architecture even though you mean product functionality, you can simply go to https://slack.com/help/articles/115003205446-Slack-plans-and...
Look at the functionality listed there and see how many things AOL supported which Slack does not and vice versa.
Seems quite the ideal use-case for Elixir/Erlang but since haven't used it myself, don't know would it be that much better. Especially training/developer pool -wise.
And Vitess/MySQL for persistence? No Cassandra or something similar?
https://elixir-lang.org/blog/2020/10/08/real-time-communicat...
Notably, with regards to the hiring pool:
> None of the chat infrastructure engineers had experience with Elixir before joining the company. They all learned it on the job.
Pretty sure YouTube migrated onto Spanner quite a while ago.
I could see these Channel Servers using something like SQLite for persistence, which would allow the full suite of SQL features without any of the drawbacks of CQL, of which there are many.
The blog post felt a bit high-level and lacking some technical/architectural details that would've given more insight, but nonetheless it was an interesting read!
Are you for real? This isn't an hypothetical codebase. This is slack, used by millions of people.
I get it, they're using boring tech, but come on.
'talk' was an everyday tool back in the 1980s and early 90s that sent characters as they were typed. I just checked and it's installed on my Mac today. It's person-to-person, not group chat, but there's something magical about seeing every keystroke in near-real-time. You can feel the person at the other end.
Its fine if you don't care enough to edit what you say, but if you do it requires composing the sentence first in your head.
Everything has gotten more sluggish.
E: if you reach out to support, they can reset stuck state for you. They were very helpful but didn't sound sure it would work.
What on earth do you even mean by that?
Talk to a controls engineer about real time operating systems and PLC programming, and you get a very solid definition. Talk to a software developer and it means something like "as fast as we can, without any purposeful delays or buffers".
Maybe I'm just a pedant, but the entire article has graphs with no numbers, no defined service level objectives, and vaguely mentions 500ms at the end.
What is real time?
No, "real time" in a software context means having a persistent stateful TCP/UDP connection open between a client and a server (or two clients in case of WebRTC and the like) to exchange data, skipping over paradigms like having to establish new HTTP connections or being blocked on database fetches. This is a well-established definition, and has nothing to do with any specific millisecond threshold.
Voice, video, and even text chat, falls into this category. You want minimal layency but nobody will die if the screen freezes for a bit.
It's customary to talk about these systems as "real-time communication" and nobody working in the field would confuse it with hard real-time.
More info and examples: https://en.m.wikipedia.org/wiki/Real-time_communication
You wouldn't ask for the deadlines of a video game speedrunner talking about "real time", either. That means real wall-clock time rather than time displayed by the video game (differences can arise due to e.g. pausing)
Maybe fixing the bugs so that they don't mysteriously get ill would be a thing to pursue?
But sure, you get a medal for being a smartass online :)
Well that's just silly, if you've ever used used slack. It's corporate chat rooms, with real-time chat being a fundamental feature of it.
> counter to what the platform is meant for.
The platform is meant for communication. Async communication is one (probably smallish, at least in my experience) slice of that.
One of the primary uses for slack, in our org, is live support. Having people sitting around, waiting for timers to expire, could never be a benefit for us.
"We have this very successful, fairly technically advanced as-real-time-as-is-practicable messaging system that most of our users use for real-time communication. We spent large amounts of engineering resources, time, and money to make it not real-time."
Above that, you have basic sharding techniques like assigning each channel to an instance. This should work up to some thousands of moderately active chatters in a single channel - at which point nobody can read all the messages anyway so it's not working for non-technical reasons.
Slack is trying to solve the problem of having thousands of channels with millions of people, as well as millions of channels with thousands of people. In other words, about 6 orders of magnitude bigger than your load test.
The per-channel limit at that rate (1 message to every users every 5 to 10 seconds) was 20K concurrent connections but the number of channels is unlimited since you can add more capacity by adding more hosts and the sharding of channels across available hosts is automatic.
The only limit was number of hosts since there is only one coordinator instance in the cluster though the cluster can continue to operate without any downtime while the coordinator instance is down. The coordinator instance only needs to be up while the cluster is scaling up or down. That said it should be able to handle up to 5k hosts, potentially even much higher.