My bet is that this incident is caused by a big release after a post-holiday "code freeze".
My bet is that this incident is caused by a big release after a post-holiday "code freeze".
- high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, the higher the divergence world of production and the world of the new feature
- lots of new hires (new year = new hiring budget). New hires are missing some tribal knowledge about the system and make a production-breaking release.
I tried to think of other reasons, but these two overwhelmingly stand out as the two biggest reasons. Would love to hear from others.
Hate on the cloud all you want, but AWS has (several flavors of) load balancers and various ways to automatically scale up and down resources (and if you're conservative, you can disable the 'down' part). If you're operating a major SaaS company like Slack and not taking advantage of them, something's gone wrong.
In 2021, how does one keep track of resource starvation at the process, container, os, service, pod, cluster, availability zone and region levels?
January is busy for recruiting, but given a week or two of interviewing and negotiating, two weeks notice, it's probably February before new employees are starting, and they're not making big, production-damaging deploys for a week or two after that.
Probably not as big of a rush as the end of school year rush in summer though.
I also doubt that new people will be breaking production on day one. Even at a fast moving startup I'd expect it to take a bit to go through the onboarding paperwork, get a laptop and actually try pushing a change to production.
People came back to work, and most of them start around the same time (US wise at least).
Hence kids - a vital lesson for all of us - don't start the call at a full hour, give it 3-7 min to make your coworkers confused and give some time for the systems to auto-scale ;)
Doubt anyone releasing big changes Monday morning.
Guess it was Slack being Slack.
> Doubt anyone releasing big changes Monday morning.
This is definitely an engineering best practice, and by best practice, I mean something that Uber's, I mean Slack's SRE team strongly pushed for, and got politely overruled on. After a code freeze is lifted, it's quite common for lots of promotion-eager engineers to release big changes.
EDIT: Guys it was a joke, chill
Releasing 1 change a year with a 100% chance of working -- no promotion for 10 years
Releasing 10 changes a year each with a 10% chance of breaking something -- 1 in 3 chance of promotion in a year, and a 2 in 3 chance of downtime
Where did that assumption come from? Also are you claiming that it takes 10x more time to release a non-breaking change?
I don't see any need to deploy a big change at once in the software world today. At worst feature gate the thing you want to do and run it in a beta environment, but still push the actual code down the pipeline.
Every Uber/ex-Uber engineer is nervously chuckling at this comment right now
Did you mean that literally? E.g. is it common at Uber that engineers can release changes to production on their own?
The engineering team is responsible for the mess caused by a bad deploy, so it's appropriate that those engineers should also choose the timing.
Our team typically deploys between 10am and 4ish, local time, since that's when we're at our desks and ready to click through the approvals and monitor the changes as they go through our pipelines.
The feature enablement happens through an EFT / beta process, and the final timing of GA enablement is a PM decision. But features are widely used by customers ahead of that time, as part of the rollout process.
Our team usually rolls out non-feature changes to services via dynamic configuration switches, so that we can get new bits in place, and then enable new behavior without a redeploy. This also enables us to roll back the dynamic config quickly if something unexpected happens.
(We generally don't do this for net new functionality; there's lower risk in adding a new REST endpoint etc. than in changing an existing query's behavior or implementation.)
Although, we also don’t close the pipeline for just any holiday break. In fact low holiday traffic is a good time to keep pipelines open, since changes will impact less people.
What it could be is some engineer somewhere coming in after the holiday, noticing a slightly flaky thing, and thinking, "I'll reboot/redeploy/refresh this thing so the flakiness doesn't get worse". Only it turns out the flaky thing was a signal of something else falling over. Or maybe the redeploy was the wrong version because of bad CI/CD, or maybe the person just fat-fingered it.
At least that how it worked at one FAANG
(I'm not making a statement if that's good or bad or if it works or whatever. Please don't read an opinion into it.)
Now, caveats etc, this was a collection of single applications in a big microservices architecture, and as the project grows it becomes more and more difficult to manage something like this, especially if you get more pull requests in the time it takes to do a build. But it is the way to go, I think.
Anyway, since tests and CI are not definitive, you also need a gradual rollout - 1%, 5%, etc - AND you need a similar process for any infrastructure change, which gets more and more tricky as you go down to the hardware level.
That sounds pretty early to think somebody on the west coast did something, other than maybe acknowledge the pages and declare the incident.
Likewise, incognito mode will also ignore most cached web content, meaning all assets on the Slack web app will get loaded again from scratch. This "clean state" start could, theoretically, get around issues with old - potentially incorrect/outdated - assets being loaded, even though that really shouldn't happen under most circumstances.