Slack’s Incident on 2-22-22
slack.engineering
slack.engineering
Many incident reports are fully lacking in any meaningful detail, or wholly unapologetic. I actually enjoyed learning tidbits about the author, in particular their mention of https://how.complexsystems.fail/.
Reading this boosted my confidence in Slack's teams, which should ultimately be the objective of a release like this. It's not pure PR nor a gruff legally-obligated disclosure.
It helps that I wasn't really affected by this incident.
From a structural standpoint I think my technical comment can be useful. If things really are failing this much A) you should figure out why and slow that down. B) if you have a generally stable system and understand the typical rate of failure, you can add tripwires into Mcrib to avoid over-culling services and loudly raise alarms. Then C) you can improve technical reliability with redundancy/extstore/etc.
I've also seen plenty of times where folks have a dependency of a service determine if that service is usable, which I disagree with quite strongly. Consul being down on a node should trigger something to consider if the service is dead. It's important both for reliability (don't kill perfectly working things because you end up having to design around it), and for maintainability as you've now made people afraid of upgrading Consul or other co-dependent services. Other similar failures are single-point-of-testing availability checking where instead you probably want two points of truth before shooting a service.
Now you risk people being afraid of upgrading probably anything, which means they will work around it, abstract it, or needlessly replace it with something they feel safer managing. The latter is at best a waste of time, at worst a time bomb until you find out what conditions this new thing breaks under.
This isn't advocating that you design without assuming anything can fail anywhere at any time; just pointing out that how often a service _should_ fail is extremely useful information when designing systems and designing fail safes, alerts, monitoring, etc.
It's a subtle difference. I think many operators get used to node failures being extremely common when they don't necessarily have to be. I suspect the note on "if they come back on their own ensure they're flushed" meaning they have something unusual causing ephemeral failures. If that's just "cloud networking" there isn't much they can do but it's almost always fixable.
And exactly how rare do you believe this to be?
In my experience, node failures at scale of hundreds to thousands of nodes are monthly to weekly, if not daily. Generally speaking, stability is a normal distribution. Young, new instances experience similar failure rates as old instances. If you have any sort of maximum node lifetime (for example, a week) or scale dynamically on a daily basis then you'll see a lot of failures.
But that generally mirrors my experience that automatic failover for stable software tends to cause more issues than it solves. A good (i.e. redundant hardware and software) Postgresql server is also so unlikely to fail that wrong detection and cascading issues from automatic failover are more likely than its actual benefits.
I'd argue that stable systems are actually worse for operational stability as you become complacent and comfortable and when shit hits the fan you are unprepared.
If they start "going bad", something is wrong. That's a signal I wouldn't want to ignore.
It has happened - once an HBA in a storage node was causing occasional corruption, another time due to a communication failure people were building things with the wrong version of something which had a memory leak and would eventually summon the OOM killer. There have been other issues.
"Have you tried turning it off and back on again" is still a terrible system management strategy.
But given the number of people I've heard using "we're on AWS, out of my control" as an excuse, this appears to be an unofficial service they offer.
I run memcached at a large scale. You are totally right. Every other year we will find ONE bad memcached node down. We use nutcraker instead of mcrouter for consistent hashing to each memcache node. Once i read "We also run a control plane for the cache tier, called Mcrib. Mcrib’s role is to generate up-to-date Mcrouter configurations" -- I was like oooooh boy, here we go....
Knowing memcache is a rock comes with experience though.
I don't believe you run it at the scale Slack does.
The people at Slack who decided to use Mcrouter (and created Mcrib) have experience running Memcached, Mcrouter and Nutcracker in production at two of the biggest web properties in the world.
Trust that they know whereof they speak.
The larger an org gets the more likely it is to do weird things to mitigate organizational difficulties be them budget, human or otherwise.
Those types of things rarely show up in postmortems for obvious reasons.
Definitely not. We host about %80 of elementary schools in the US. Not slack scale but definitely face many of the same issues :/
Across the whole fleet (all services), we lose 1-10 servers per day as a baseline. Major events are then on top of that and can impact thousand of hosts at once.
Granted, this is easy to say once the incident happened with an excellent postmortem, but this should be an industry-wide wakeup call: don't do this.
I have the same issue at work, where people treat a "prometheus node_exporter down" as a "the app on the machine is down". I've started to add the actual app name in our alerts, and now people don't freak out anymore when they see "down" alerts: oh node_exporter is down, but not the app? Don't panic and calmly check why.
Also caveat: I have no deep view into Slack's infrastructure so anything I say here may not even be relevant. YMMV.
First some self promotion: https://github.com/memcached/memcached/wiki/Proxy memcached itself is shipping router/proxy software. Mcrouter is difficult to manage and unsupported. This proxy is community developed, more flexible, likely faster, and will support more native features of memcached. We're currently in a stabilization round ensuring it won't eat pets but all of the basic features have been in for a while. Documentation and example libraries are still needed but community feedback help speed those up tremendously (or any kind of question/help request).
It's not clear to me why memcached is being managed like this; mcrouter seems to only be used to abstract the configuration from the clients. It has a lot of features for redundant pools and so on. Especially with what sounds like globally immutable data and the threat of cascading failures during rolling upgrades it sounds like it would be very helpful here.
If cost or pool sizes are the main reasons why the structure is flat, using Extstore (https://github.com/memcached/memcached/wiki/Extstore) can likely help. Even if object value sizes are in the realm of 500 bytes, using flash storage can still greatly reduce the amount of RAM necessary or reduce the pool size (granted the network can still keep up) with nearly identical performance. Extstore takes a lot of tradeoffs (ie; keeping keys in RAM) to ensure most operations don't actually write to flash or double-read. Extstore's in use in tons of places and everyone's immediately addicted.
Finally, the Meta Protocol (https://github.com/memcached/memcached/wiki/MetaCommands) can help with stampeding herds to help keep DB load from exploding without adding excess network roundtrips under normal conditions. I've seen lots of workarounds people build but this protocol extension gives a lot of flexibility you can use to help survive degraded states: anti-stampeding herd, serve-stale, better counter semantics, and so on.
Any data that is fronted by the memcached caching tier can tolerate staleness right? There is not much difference between a short TTL of 1s versus an async replication delay of 1s.
Abusing vitess primaries is the root cause of this incident, and a similar incident can happen even without any scatter query.
(The original McDonald's McRib sandwich is well known for only being sold a limited time.)
So "Mcrouter" comes from Memcache-Router, then the obvious McDonalds jokes are made and someone cleverly suggests "Mcrib" for the next service. But I can't think what the backronym would be for it. Memcache Ring Buffer maybe. Or Broker.
P.S.
Yes I know the uptime is decide by committee, and doesn't reflect reality. I am just being cynical.
https://www.jagranjosh.com/general-knowledge/22-02-2022-is-b...
At the same time it seems like a horrifying job I would never ever want :D
Even the largest Slack instance probably has under 100,000 users and less than 1000 peak messages per second. That feels to me like it could be served by a single master DB. It feels like that would be a better way to shard.
Major downside is the major difference in shard sizes so some management/migration might be needed but it seems doable to me.
Certainly it feels naively that scaling should be easy due to the way slack instances are independent (unlike say Twitter).
(I'm also surprised/reminded how fast Slack grew and how quickly it became effectively ubiquitous — I think every company I've contracted for in the last five years has used Slack).
This is not true, by an order of magnitude.
I am really happy you think so. But no. We are really not transparent about them. At all.
There is no 22nd month, so we know the 22s are the day and the year, leaving only the 2 to be the month. Is it really that difficult to parse?
It shouldn't be this hard.
Think about it like speaking a different language, except with numbers and not words.
Americans memorize inches and yards, and often also memorize centimeters and meters, and working with either is fine, but we're not so often faced with numbers where it might be inches or centimeters and we have to figure out which (and when we are, it's sometimes a pain - certainly a bigger pain that working with known units).
Or, working with your language analogy, please go fetch me some "pasta" without knowing whether I'm speaking Italian or Polish.
… all the major predominantly English speaking counties will use mostly hyphens in the dd-mm-yyyy format. So although there is ambiguity, it’s easily resolved by picking that as the default mentally and only back tracking on failure.
Now in the more general case, this whole thing feels like a lieutenant/leftenant situation. We are annoyed simply because it’s not they way that we do things in a peculiar case, when otherwise the language is fully intelligible.
Picking one default and back-tracking on failure really isn't that comforting nor the constant reminder that the date you thought it was might be something else.
Yes, which makes reasonable the complaint about mm-dd-yyy.
> So although there is ambiguity, it’s easily resolved by picking that as the default mentally and only back tracking on failure.
This is both more work and also error prone in the general case (although it works out fine in this case).
> Now in the more general case, this whole thing feels like a lieutenant/leftenant situation.
Not at all. Whether I read "lieutenant" or "leftenant", I know what you're talking about. If I read 2-10-23, I might miss your birthday party.
The correct analogy is I don't know which language is spoken and the same words get used in multiple languages with different meaning. Now I can apply heuristics to figure it out or in some cases I can only guess.
I think you meant "to be the month" there. qed■
What do I win?
None of their other incident reports even have a date in the title. Yet this one does, and in a weird format. Maybe there's something novel about the date, and it's written this was to emphasize the novelty, not to provide some vital information that happens to be excluded on every other incident report title they've posted.
a) Use a far superior date format which nearly the entire world uses by default and is better and simpler in many ways.
b) Do logic when we see dates to try workout what format the date is in.
Going with a seems like a no brainer...
22 Feb 22
Feb 22 22 (weird but still better)
22 22 Feb (very weird but still better)
This also goes to show that 2022 is a better choice. My own personal preference - the 22nd of February 2022.
There is a reasonable argument for little endian dates (as in the least significant information is usually the most relevant as it changes most often), but apart from the "it has been like this forever" I don't see any reasonable argument for middle endian date formats. Then again, the US is notoriously resistant to the metric system too.
All that said, I definitely agree with the original complaint, m-dd-yy is an atrocious format. If you're going to use dashes, stick with yyyy-mm-dd. Replacing the dashes with slashes, as in 2/22/22, would have been fine.
People I have met from Australia, South Africa, the UK, all have the same flexibility.
Edit: [1] says my hypothesis is most likely wrong, but that the UK just changed it later to match the rest of Europe. So maybe that influenced their way of speaking? In any case, matching the way one speaks doesn't seem to be a strong reason as it's easily adaptable and month names are unambiguous. Interestingly it also quotes that using a purely numeric format is incorrect in any formal use as to not confuse month and day.
[1] https://iso.mit.edu/americanisms/date-format-in-the-united-s...
:)
Putting the day first isn’t actually a benefit because you still need the year and month for context.
Any favoritism of the EU format is the same as the US format. Just familiarity.
ISO8601 is the only way.
Since you are already on the fence with ISO8601 I invite you to consider time of day. Would you use second:minute:hour? That is also in (reverse!) “chronological magnitude” order.
If you want to write the date little endian then you should do the same with the year. So today’s little-endian date is 26-04-2220. Or maybe that is 62-40-2220? Or is it 62-40-2202?
ISO8601 is the only sane date format. Anything else is only favored for familiarity.
In many other English-speaking countries people usually say "the second of March, nineteen sixty two."
twelve thirty or half twelve
color or colour
And endiannes / sorting comes up in real life pretty often - scanning for large numbers in the price list, or finding stuff in the sorts list.
I think if history turned differently, we could have had sane time format in the US.
If you want to write it out the way it’s spoken, write it out the way its spoken. Mixing the computer’s numbers and the spoken word’s grammar make for misunderstandings, and as a programmer, eliminating misunderstandings is one of my goals.
If I'm writing a letter or addressing a specific thing in a formal context I chose to "revert" to the Month, Day Year because it's the social standard for the country I am in, and I want to fit into that cultural expectation, but if it's for a business document or normal chatter I think DD MMM YYYY is probably the clearest to both English & Non-English speakers. It eliminates the distractions I'd normally be dealing with when considering if I'm talking to someone out of country or not. It would be really great if it ends up being more widely adopted.
The only time I've every heard someone say "<Month> <Ordinal>" or "<Month> the <Ordinal>" is when talking with Americans.
Every other time it's always "<Ordinal> of <Month>" or just "<Ordinal>" for short.
It doesn't seem that bad to turn it into 2-22-22.
In this case you can lookup Slack outages to disambiguate it, but the frustration here - and I share it - is directed at the stubborn refusal to use a standard format that the reest of the world has agreed upon.
Yes, the numbers are all the same, and the author is based in the US, and thus is using the default format in the US. So odd that this is the top comment.
(i had scrolled immediately down, so the thread titel wasnt visible when I was reading your comment)
haha
Not sure if this is standard but I usually see the delimiter being used to define the date format: big-endian y-m-d uses dashes, middle-endian m/d/y uses slashes, and d.m.y little-endian uses dots.
No. Pretty much every separator / order combination is in regular use.
See the table under “Listing” at https://en.m.wikipedia.org/wiki/Date_format_by_country
a) m-dd-yy
b) m-yy-dd
I can't tell. How do you disambiguate?I agree with you though, the point of a date like yyyy-mm-dd is to avoid working out stuff like this. You don't pick a date format based on whether the current date is ambiguous or not.
Agreed, this is why ISO8601 exists.
It's the title because it's a novel date, and formatted this way to emphasize the novelty. It's also a date for which your question is irrelevant: it's the same either way.
It’s the ordering that’s significant, not the separator.
A mixture is even more likely to be d-m-y, today is 27/4-2022 in Danish handwriting.
I've never seen a date written like this, interesting.
Yeah, I guess if people look at it and parse it, they understand it. But what I notice is, that the IT-bubble I am in has no worries parsing these dates because we use them a lot in IT, even in Europe. But people outside of IT do seem confused about US-formatted dates from time to time, because they rarely encounter them.
More than once did I notice someone struggling to fill in their birthday into an online form because the people making the form decided to use M/D/Y instead of D.M.Y Of course most can help themselves, but its not like these date formats seem natural or normal to everyone in Europe.
2-22-22 was also when Russia invaded Ukraine
And Joe Bidens statement about the invasion was on 2-22-22 2:22pm on the dot.
I could not figure out the significance of this more than 11:11 was when WW1 ended, but it's probably something else.
Then Biden did a speech against on 2:22 22-2-22
The actual war started in 2014
https://edition.cnn.com/europe/live-news/ukraine-russia-news...
Old man yells at cloud vibes here.
> It's also a data breach incident (denial of service, unavailability)
How is unavailability a data breach?!
TLDR;
https://www.lexology.com/library/detail.aspx?g=03e8a988-7c9e...
As if no other way to communicate exists?
I remember using Slack, feeling fed up with emails, until I realized that if I wanted to sync Slack messages offline and have a standard way to view these messages that I was SOL. I am so glad that I've returned to email and optimized my workflow to use email effectively and efficiently. The best part is no more vendor lock-in.