Slack’s migration to a cellular architecture
slack.engineering
slack.engineering
A past team of mine managed services in a similar fashion. We had a couple (usually 2-4) single AZ clusters with a thin (Envoy) layer to balance traffic between clusters.
We could detect incidents in a single cluster by comparing metrics across clusters. Mitigation was easy, we could drain a cluster in under a minute, redirecting traffic to the other ones. Most traffic was intra AZ, so it was fast and there was no cross-AZ networking fees.
The downside is that most services were running in several clusters, so there was redundancy in compute, caches, etc.
When we talked to people outside the company, e.g. solution architects from our cloud provider, they would be surprised at our architecture and immediately suggest multi-region clusters. I would joke that our single AZ clusters were a feature, not a bug.
Nice to see other folks having success with a similar architecture!
We have ~150 accounts, so roughly 75 different department, with some having not much and others have a lot of resources.
Its complex, but provides a lot of nice security primitives. We have an overarching administrative account, but that doesnt get used (and lots of alarm bells go off when it is).
I'm mentioning this for other people to be aware as one can easily make the assumption that an AZ is the same concept on all clouds, which is not true and painful to realise.
Clearly define boundaries and acceptable behavior within boundaries.
Setup up telemetry and observability to monitor for threshold violations.
Simple. Right?
fewer things to monitor. fewer things that can fail.
just log in from time to time to update packages.
see, cloud doesn’t have to be complex.
sure, use "pet" computers for experiments and dev.. but having produluction be a "cattle" makes your life so much less stressful.
E.g. you want to think about your machines/containers like cattle, so you put them into a kubernetes cluster, which has become your new pet. If all your infra fits on one machine, it's way easier to have that as your pet the same way it's easier to have a dog than run a livestock operation.
> If all your infra fits on one machine, it's way easier to have that as your pet
What stops you from automating the provisioning of that?
These are two orthogonal issues. Clusters are used to manage many machines. If you need only one machine, you don't need a cluster. Either way, the provisioning of single nodes and clusters should be automated. Whether you have pets or cattle is not related to how many machines you have.
(1) Raw scale might mean you just plain can't fit everything on a single t2.2xlarge.
(2) Different services (containers) might have different performance profiles, so you may want a few different types of machines around.
(3) You probably still want N+2 redundancy even within your single AZ, so this scheme should at least be upgraded to three t2.2xlarge boxes. ;)
I’m an EE by education. It’s electron state, silly leaky abstraction Stan’ing to my head.
Different babble for allocation of memory and algorithmic manipulation of the values stored within.
Correctness is important when it comes to results being mapped to human consumption and even then the subset of parameters to be be rigorous with can be made subj. personally I lean on a subset that includes biological health and well being and deprioritize religiosity
I did the math for our own stack and after a setback month in client revenue, and decided to put all our servers into a single AZ in a single region. The only multi-AZ, multi-region services are our backups. Surviving bad machines happens often enough that it's priced in via using Kubernetes, but losing a whole AZ is a freak accident that's just SO rare that calculating real business risk, it seemed apt to pretend it just doesn't happen (sorry, Google Cloud Paris customers).
Call me reckless, but I haven't looked back ever since and it saves us thousands of dollars in intra-AZ fees per month alone.
People building this huge Multi AZ, Hyper Redundant, Multi-National, infinitly scaling Cloud Solution for something that requires a single VM and a Database.
Most Companies just don't need that level of scale and would be better off building something smaller and when you actually do scale you rewrite it with the profits made from the smaller solutions.
Of course there are many companies that do require something large but you should seriously consider if something smaller will do first.
I think solutions like a 100% cloudflare workers based backend can sidestep this a little but usually, that's not possibly or even the right thing in every situation.
Most of the situations where we needed to drastically scale up were known ahead of time as well (e.g. campaign from customer), and we would preallocate instances or even more clusters.
I may be forcing my memory, but if I'm not mistaken, our auto scaling was setup in a way that the system could handle sudden load increases of ~50% without noticeable disruption. Spikes bigger than this could lead to increased latency and/or error rate.
That said, it's a trade off between efficiency and load spike tolerance. I trust that the trade off is made with informed decision.
Unless you co-mingle online and offline (batch) traffic on same hosts, flat response times and high utilization aren’t compatible.
I don't think that relatively low utilization rates is the scenario that requires "informed decision". The only tradeoff in low utilization rate scenarios is cost, which might be outright cheaper and irrelevant once you do the math on the tradeoffs of using reserved instances vs the cost of scaling up with on-demand instances.
You need to make a damn good case to chronically underprovision your system and expect it to autoscale your way into nickle-and-dime savings.
Worth noting that requiring teams to use 3 AZs is a good idea because you get "n" shaped patterns instead of mirror shaped patterns, which have very different characteristics for resilience and continuity.
AWS does that with their lambda arch to reduce waste.
Not being a dick here but is this not a fairly obvious flaw?
I mean why not keep a structured "message log" of all channels of all time ?
For every write the system updates the message log.
I am guessing and making assumptions I know.
Having worked on other messaging apps, these are usually separated because their performance/scalability requirements are different
Agree with this point of view. Except the Jabber/XMPP Cisco legal thing, there's just no tech answer on why on earth Slack did not use XMPP under the hood.
Slack would have known about WhatsApp architecture because it was widely talked about pre-FB acquisition (2014).
And Slack was founded in 2013.
Sometimes it's for OKRs and power/politics balance between departments and teams in an enterprise. "If an existing tech like XMPP is used then no serious development could be needed" fear (which is not really true). It can lead (and leads) to a huge waste of resources and overspending.
But it's a bit similar to building a luxury house. Not because an owner needs it. But because he can afford it.
We had it running accross bare metal, AWS and Azure and it one of the key aspects was that it handled persistent workloads for big data, including distributed databases.
Kubernetes was just getting built when we started and was supposed to be a Mesos scheduler initially.
I assumed Kubernetes would get all the pieces in and make things easier, but I still miss the whole paradigm we had almost 10 years ago.
This is retro now :)
Wouldn't that be the same thing, with the obvious caveat you are t using the routing technology Slack is using (We don't - We use vanilla AWS offerings)
[0] There are of course exceptions that come once every few years, but most instances people can think of in terms of widespread outages is one specific service going down in a region, creating a cascade of other dependencies. e.g. Lambda or Kinesis going down and impacting some other higher-level service, say, Translate.
Not at AWS: https://aws.amazon.com/about-aws/global-infrastructure/regio...
> An Availability Zone (AZ) is one or more discrete data centers with redundant power, networking, and connectivity in an AWS Region. AZs give customers the ability to operate production applications and databases that are more highly available, fault tolerant, and scalable than would be possible from a single data center. All AZs in an AWS Region are interconnected with high-bandwidth, low-latency networking, over fully redundant, dedicated metro fiber providing high-throughput, low-latency networking between AZs. All traffic between AZs is encrypted. The network performance is sufficient to accomplish synchronous replication between AZs. AZs make partitioning applications for high availability easy. If an application is partitioned across AZs, companies are better isolated and protected from issues such as power outages, lightning strikes, tornadoes, earthquakes, and more. AZs are physically separated by a meaningful distance, many kilometers, from any other AZ, although all are within 100 km (60 miles) of each other.
This is unique compared to Microsoft, and Google (a single flood taking out multiple AZ's? Uh oh: https://www.theregister.com/2023/04/26/google_cloud_outage/)
Sure, a massive earthquake or a nuclear strike could probably take out several.
Which is clearly false if a single flood can take out an entire region.
Azure doesn't guarantee distance between zones, but they aim for 483km. So they can provide better isolation at the cost of higher inter-AZ latency. Depends on the region - you'd need an internal contact (and/or NDA?) to get approx numbers
https://www.google.com/search?q=us-east-1+reliability https://www.google.com/search?q=us-east-1+outage
us-east-1 does have more problems than other zones due to a variety of reasons, but it rarely (ie, once a few years) goes down as a whole. As long as you're in several AZs within us-east-1, the impact of most outages should not take you down completely. In the context of the comment you are replying to, your google search links are lazy and fail to see the big picture.
(p.s. the use of a google search vs direct results wasn't "lazy" - it's to allow readers to do their own research vs pasting one result and then getting accused of bias)
edit: Hmm or maybe not? I still sometimes confuse aws terminology. Perhaps it is all in us-east—1, just in different availability zones (buildings?)
Correct, us-east-1 has several AZs, names like us-east-1a, us-east-1b etc. IIRC us-east-1 has six of them now.
So if one goes down nothing is lost, but capacity and durability is degraded.
Or is this some lower-level service that “doesn’t count” somehow?
You always need cross-AZ traffic, otherwise your data is single homed (which we used to call "your data doesn't exist").
Also explains to me why new features would take a while to roll out of you are cautiously updating instances/AZs one by one
We use a few languages to serve client requests, but by far the biggest codebase is written in Hack, which runs inside an interpreter called HHVM that’s also used at Facebook.
Plus, it seems telling that Threads was developed in Python - not Hack.
(I’m aware IG is Python & it’s the same team)
> PHP makes it really easy to make a dynamically-rendered website. PHP also makes it really easy to create an utterly insecure dynamically-rendered website.
As a related thought, a lot of the modern serverless stuff feels like it's reinventing the ideas of Apache + PHP, or perhaps CGI?
Thanks for Psalm!
Curious, if Slack was built today from ground up - what tech stack do you think should/would be used?
A slightly different question that’s a bit easier to answer: “if I could wave a magic wand and X million lines of code were instantly rewritten and all developers were instantly trained on that language”.
There the choice would be limited to languages that have similar or faster perf characteristics to Hack, without sacrificing developer productivity.
Rust is out of the question (compile times for hundreds of devs would instantly sap productivity). PHP, Ruby, Node and Python are too slow — for the moment at least.
So it would be either Hack or Go. I don’t know enough about JVM languages to know whether they would be a good fit.
Some follow-up …
A. isn’t PHP on par perf wise to Hack these days? Re: “PHP is too slow” comment.
B. have you ever looked into PHP-NGX? It’s perf looks impressive, though you lose the benefit of stateless
No. But I don't have any numbers, because it's been years since the two languages were directly comparable on anything but a teeny tiny example program.
Facebook gets big cost savings from a 1% improvement in performance, so they make sure that performance is as good as it can possibly be. They have a team of engineers working on the problem.
PHP doesn't have any engineers working on performance full-time — it's impossible for the language to compete there. Hack has also removed a bunch of PHP constructs (e.g. magic methods) that are a drain on performance, so there's no way to close the gap.
But that should in no way make you choose Hack over PHP. Apart from anything else, the delta won't matter for 99.9% of websites.
Other alternatives are Apache, Caddy, and more…
>Slack does not share a common codebase or even runtime; services in the user-facing request path are written in Hack, Go, Java, and C++.
… but it’s probably JSON and some JSON-Schema-based “now you have two problems” junk instead of what I described. In which case, yeah, ew, gross. Unless they’ve made some unusually good choices.
If no new requests from users are arriving in a siloed AZ, internal services in that AZ will naturally quiesce as they have no new work to do.
Not necessarily because, due to some bug, there may be resource-hungry jobs running indefinitely. (Slack's engineers must have considered this; I am just nitpicking this particular part of the text.)Cell-based architecture introduces modular 'cells' into software systems, each with distinct APIs. This design fosters loose coupling and scalability – key for today's dynamic software landscape. Particularly, for those intrigued by microservices, cells align seamlessly with the independent, scalable components that power microservices architectures.
Curious to dive deeper? If you're keen to explore the nitty-gritty technicalities, I invite you to check out the architecture paper https://github.com/wso2/reference-architecture/blob/master/r... for an in-depth understanding. Let's kick-start this dialogue on the potential of cell-based architecture and its impact on modern software design. Feel free to join the conversation!
However, this article uses "cell" in a completely different way. It is not the cell-based architecture that you are promoting here without reading the article.
Semi-related tangent: sometime around mid-2016, I came across a tool that helped visualize requests in near real-time, and showed what it "looks" like (ie, flow slows to trickle in service A during draining, while it ramps up in service B)... there was a really compelling demo, but I never bookmarked it and can't seem to find it. IIRC its name was a single word. Maybe someone reading this will know what I'm talking about... ?
> If a graph of nodes and edges with data about traffic volume is provided, it will render a traffic graph animating the connection volume between nodes.
How would one go about providing such a graph? :)
This kind of shared communal knowledge is one of many reasons I'm very grateful for the HN community.
> This turns out to have a lot of complexity lurking within. Slack does not share a common codebase or even runtime; services in the user-facing request path are written in Hack, Go, Java, and C++. This would necessitate a separate implementation in each language.
This sounds crazy. I've seen several products where there is a core stack (e.g. Java) and then surrounding tools, analytics etc in Python, R and others. But why would you create such a mess for your primary user request path?
Sure, they're not "just a chat app" they have video, file sharing etc included and a lot of integrations. But still this sounds like a company that had too much money and too little sense while growing rapidly.
They see the job as a strictly technical problem looking for the best technical solution. They don't look up and see how that problem fits into the larger organization.
They think things like "I can make a microservice that encodes PDFs 10x faster by using Rust" and give an estimate based on that, never thinking about how we're going to need to hire 2 more Rust devs to keep that running, and we could have delivered twice as quickly if I had used our default Python stack and now our "10x faster" doesn't matter because that feature is old news.
Microservices are such an unfortunate concept because they attract the people least suited to use them: If your team can't handle a monolith, you shouldn't even be looking up what a microservice is.
It's a sad state of the world that almost every application now is written in Javascript and deployed with Electron, and massive memory usage and slow UIs have become accepted as the norm.
Try any IRC client and tell me, with a straight face, that Slack is just as responsive.
Slack won’t seem that bad, until you use something that’s actually good.
The loss of performance on commonplace applications has been a real “boil the frog” situation; we’ve lost so much performance and responsiveness, but it’s happened so gradually that most people don’t notice.
For large scale, cross cutting initiatives you’ll have some pain. For feature velocity, you’ll see great results. Everything is a trade off.
Even the percentage of nerds that would want IRC or XMPP bridges back would have to be vanishingly small. I’d be annoyed if Slack reimplemented such functionality because it no doubt slows down future development. Slack has a number of mechanics that do not carry across to IRC or XMPP, and they did when they killed the bridges. I’d be annoyed if new features were compromised to increase compatibility with this blatant nerd vanity project.
I would be fine with the understanding that the IRC bridge was missing functionality (and it always was). Although threads might make it impossible to implement in a nice way now.
As far as new features go, I don't want any new features in Slack: it worked exactly like I wanted it to seven years ago and the new stuff is nice, but not worth the degradation in user experience.
Slack IRC bridging in the 2014/2015 era was great. We had a lot of people who spent their whole workday in a terminal window and weren't interested in running a web browser in the background continuously just for a chat room.
Yeah, although one can dream that some SaaS company would do htings differently
No! Stop diluting this word.
Yes, you're right, I'm misusing it.
However, I think that there is a phenomenon that happens to a lot of tech products that is more general than what Doctorow is talking about. There is a certain type of person who is attracted to building a new thing, and there is a different type of person who is attracted to a thing that is already successful. Pioneers and Settlers, as a former colleague of mine described it. In the context of Internet services, pioneers care a lot about attracting users initially so they tend to dwell on every minor detail. Settlers care a lot about stability, so gradual degradation over time (e.g., in performance, in other measures of quality) is tolerable as long as its rate is controllable and well-understood.
I think that Doctorow's thesis is a special case of this where greed is the driving factor behind the gradual erosion of quality.
(Or, now that I notice your username, maybe you’re making an ironic joke, since complaining about the misuse of the word enshitification is a meme now?)
Not saying they shouldn’t fix their reliability. Every other week it seems like they have an outage with this or that.
The Flickr style commit to production multiple times per day seems to have its limits. Perhaps longer canary and slower rollouts would help.
If they had "good infra people" then their program wouldn't sit and spin for seconds, ever.
What? Does amazon need to push for new sales points or are they simply making up architectures now?
http://highscalability.com/blog/2012/5/9/cell-architectures....
But their implementation isn’t as grim as what I had initially envisioned when hearing that term. I immediately thought of Smalltalk and the idea of objects sitting next to each other, forming a graph (of no particular structure… just a graph), passing messages to neighbours. Like cells in an organism send hormones and whatnot. That makes for a huge mess that cannot be reasoned about, hence why we instead went with stricter structures like trees for (single) inheritance. That’s much closer to this silo approach, which seems nice and reasonable (although I get the impression considerable complexity was swept under the rug, like global DB consistence; the siloes cannot truly be siloed).
This blog is writing about availability zones as if they're a new concept too.
May be marketing but it is an architecture born out of Amazon's (and AWS's) use of AWS:
- Reliable scalability: How Amazon.com scales in the cloud, https://www.youtube.com/watch?v=QeW9wCB36ck&t=993 (2022)
- How AWS minimizes the blast radius of failures, https://youtu.be/swQbA4zub20 (2018)
For massive enterprise products like Slack that need close to 100% uptime across all their services, cells make sense.
Here's AWS's 2019 guide for financial services in AWS, where the isolated stack concept is referenced under parallel resiliency section and called "shared nothing":
https://d1.awsstatic.com/Financial%20Services/Resilient%20Ap...
"There are three dominent themes in building high transaction rate multiprocessor systems, namely shared memory (e.g. Synapse, IBM/AP configurations), shared disk (e.g. VAX/cluster, any multi-ported disk system), and shared nothing (e.g. Tandem, Tolerant). This paper argues that shared nothing is the pre- ferred approach."
It's a blunt tool, much like PHP. PHP does seem to be a good choice for them, but I wouldn't want to work there. It's all right, there are different ways to do stuff.
At the time, we had healthy debates about whether the feature was useful enough to justify additional complexity, and whether there could be cases where it would backfire. To this day, it's an underused feature. I still regularly run into customers and configurations that cause unnecessary blips to their end-users, so it's nice to see when people dig in and make sure that the next level of networking is working as well as it can.
[0] https://news.microsoft.com/1998/08/24/microsoft-corp-acquire...
No reserved capacity (pay for usage), so it works for boot strapping startups and provides superior resilience while being extremely simple to setup and involves almost zero maintenance or patching (even under the hug of death). I don’t understand settling for less (and taking longer and paying more for it).
This writeup seems like a useful contribution to spreading this knowledge that you think every engineer should, somehow, innately be born with, to those members of the development community who missed out on picking this stuff up in elementary school.
Sorry - where are you quoting this claim from?
But yeah any good SRE could point this out years ago.
Teams especially, is something I loathe using every day. Everything about the UI and UX gets in the way of what I’m trying to do, rather than assisting in or even enabling it. It’s like it doesn’t want me to communicate - it wants me to react and offer as little useful information as possible.
Eventually the list will grow so large that we could probably attach a 5-figure dollar amount to it, if it hasn't already.
Speaking of which, I'm going now to buy more Microsoft shares.
I don't think anyone was sad that slack didn't integrate with the other MS services "stack".
(Teams still can't copy images, instead you get a massive base64 block of text iirc)
Honestly, bugged me a few times before I just switched to using the snippet tool. I use it all the time anyways and this felt natural.
Does it really though? In my experience teams has a buggy integration with other things in the stack.
And Teams itself ia massively buggy and a resource hog for the whole time I've used it.
Since then, they have hit 2 home runs on top of their basic chat functionality:
1. Slack Connect: Being able to share channels between workspaces is simply amazing. Most of our customers are on Slack, and having a Slack connection to them makes it much easier to communicate with them and get their feedback as we improve our product. I don't know any other tool that even comes close to how important Slack Connect has been to product development in my startup.
2. Canvas: They rolled this out last year as notes or something, and I was pretty underwhelmed with the experience at first. But very recently (within the past month I think) they reintroduced this as "canvas" with really tight integration with threads. We have moved all of our planning and synchronization activity to a canvas that we set up every week.
Although these features are not difficult for Google or Microsoft to implement on a purely technical level, their product organizations don't seem to understand the network effects of chat the way that Slack's product organization does.
Slack is certainly not dead today and they are showing the savvy to stay alive well into the future.
And besides that, the UX of Teams is miles behind Slack.
I haven’t regularly used teams in about a year, but I would legitimately consider passing on a job offer where they used it.
In a thread where many folks are talking about using the best tools for a job, teams is never the best tool for any form of digital communication.