We ship the most code on Friday
blog.railway.app
blog.railway.app
Why hire any QA? Why code review at all? All of these reduce the risk of bad or unintended software being in the wild and not releasing on friday or after 4pm reduce the amount of time it can be in the wild without harassing people who are otherwise on their off time. You can have a half dozen full time QA people and a horde of reviewers merging in your request into a staging instance that triggers E2E tests that complete and require a manager to merge into master to be automatically deployed via the same docker image into a cluster and the only guarantees you have is that it may be absolutely correct and that it may have an issue.
From a more pessimistic perspective, as times goes -- code will likely get more messy and you'll likely have more customers (assumings things are going good), in which case the fact incidents are not happening now, wont mean it wont start happening eventually.
By all means deploy to staging on Friday and have some automated and load testing running on weekends if that's what you want to do. But there's really no reason that promotion to prod should happen in a Friday or just before any significant time off.
It really doesn't have much of a justification. You want to "feel good" about your work week and ship on a Friday? Ok great. How is that benefiting your customers? It's not urgent, it can wait until Monday or what have you.
There are very few hard rules with engineering.
Not understanding that your internal users are your customers is a mistake.
But it's also a failure when being slow with tech due to fear of a few bugs can hold back the entire company. Some bugs are showstoppers (bring down the system type of things). Others are very inconsequential. It's up to us as engineers to figure out how to use that context to avoid the first type, but make informed decisions about the second type. To be honest in 15 years of working with my current teams we have never had a bring the system down type of bug. It just doesn't happen.
If you're not confident about deploying on a Friday, you shouldn't be confident deploying at all.
To be clear, for big, risky changes, we obviously waited for the following Monday. The key is we didn't have a blanket ban on Friday deployments. Not being able to deploy for 20% of the working week is bad.
If you are so confident why even have an oncall? Because we know reality is. Bugs are rare and severe bugs are more rare. But they do happen. So let's try to make them happen less when people are out of office.
Are you upper management at my job?
Deploying when you have low usage has a lot of benefits. But the article is convinced that the difference is really because they are much better than anybody else.
Btw, 102 incidents within a week. Whoa! That's a lot. Wouldn't want to work there. To put it into perspective, here at AWS my team has about 3 incidents within a week.
I might be wrong but I'm pretty sure that's every incident they've ever documented, grouped by day, as opposed to the data from last week
Their status page seems to suggest ~99.97% uptime (about 3h/year down) recently
If your teammates are deploying risky changes and purposely doing it when you are outside of office hours I think that is quite disrespectful. For most teams oncalls are for emergencies, they aren't paid for full-time work and they likely have better things to do with their time than come in to work to clean up your change. Unless you are coordinating before time I would avoid doing this.
Exactly. But you get to decide how to spend your own weekend.
Depends on the exact escalation instructions, of course. Once upon a time I could be reasonably sure that if anything breaks in production and I'm on call, no one else would get called.
Yes, exactly! Don't ship after 4 PM either!
Things can and do break. It's impossible to test everything. That's why.
Your wording seems to indicate that you work by yourself. That's fine, but things change when large teams and systems are involved. You can't vouch for every single change or the inter-dependencies between all systems.
It's pretty elitist too. "Your code only breaks because you suck" doesn't make for good software engineering practices. Everyone makes mistakes.
If you have a micro service that takes requests over / foo and /bar and returns JSON that needs to fit certain properties, and you’re submitting a change to fix a corner case bug or add a new property that isn’t read, sure.
If you are working on low level networking infrastructure that is used in tandem with thousands of different workloads you need to support, won’t work.
Eventually, your totally avoidable gamble will backfire.
You know where I did get pings all the time? In companies that had strict release windows and would never deploy on Friday. Which in result motivated people to "push now & hotfix later" to meat the release cutoff date.
For people who work on applications they are often only hurting themselves if things break, so the risk profile is different. Maybe you don’t even have any customers, maybe your customers are consumers who are just not able to place orders.
IMO these groups often talk past each other (although as someone who works on infrastructure I think we often consider application developers’ more than they consider us, because they are our customers, and to them we are doing best when they don’t need to think about us) which leads to all these arguments from “Ship fast ship often - nobody uses our mostly static website which is easy to test, so who gives a fuck if it breaks” to “canary every change and carefully A/B test it with long release times to ensure it is safe”.
And fair on OT for weekend work, but again, a lot depends on what was done to fix. If it was a rollback and retry next week after fixing any findings, this is a lot less bad than the high energy frustration of trying to rollback during peak hours.
I recommend that you start with "anytime as in normal working hours", which means having load-balancing/failover in production so that you are mildly degraded but no real outage. A lot of shops never get to this point and you'll find that it's awesome compared to the old "just stay up late and do it off-hours" routine that creates even more problems. The thing I like most is the ease of doing incremental daily rollouts instead of big-banging the shit out of everything at once because staying up late every night for a week is unrealistic.
Maybe someday you'll get to the ideal peak of everything-automatic-regression-testing that is unrealistic for most of us, but normal working hours is a pretty good second best.
If you do a 7 day rollout on Monday then you’ll be hitting those huge chunks of users over the weekend which may run against rollout strategy. This is to say nothing of review times… which have been decently fast the last few years.
But really, as a solo developer I ship on Fridays because it coincides with a weekend where I’ll have time to deal with issues in a low stress way (not as busy).
My situation is somewhat hybrid because I publish to MAS and direct to my app’s website.
minimize is the keyword here. I assume you have seen regressions with Friday pushes and team had to spend their weekends resolving that at least once.
People will do mistakes, bugs will go into code, regressions will happen. It's just inevitable and no amount of automation can cover all your user paths.
People will hate this unpredictability to plan their weekends and spend a peaceful weekend. IMO, it's not worth it.
But there is one important point I'm think they have learned but they do not explicit say in this blog post.
It's what you deploy. Context matters. Deploy most things as normal on Friday's.
But you do hold off the release that is bigger than normal and involves some new infra on the backend side or is known to more memory heavy etc.
But for bug fixes etc. Deploy. Works great here at least. And I am, or at least for most of my career, was the person that fix the problem.
You ship it you fix it. Then it's up to the person that has the PR. A judgment call
If so, unless they always squash when merging, this just shows that they have the most commits on Fridays, not deployments.
edit: I updated the image caption to specify. Thanks for pointing that out!
This kind of deployment also brings more discipline to the process, because you have to do the switchover when users are actually using the system. It means people are a lot more conscientious about making sure their plans actually work.
We try to push on Monday afternoons, but it generally turns into Monday night.
At some point your app doesn't/can't really have downtime, which is also a factor.
Yes, the shipping process must be automated, tested, have safeguards (acceptance tests and an automated rollback process). But it doesn't justify taking a risk of ruining a weekend for your entire team.
Over time, working remotely, many of my team are in different time-zones. Some in London, New York, and L.A. which all have wildly different time-zones. I have to be sometimes like a cat (who are known to sleep a lot since they need to build up enough energy to pounce on wild animals without notice and do that straight away without hesitation).
So I'm used to being paged in real-time to 'pounce' on critical issues. I'm on-call like a fireman. If anyone's working remotely here reading this, this is a skill that is learned.
I make my team stay off Slack on weekends. Unless there is major incident. So would not ship on Friday to risk having them fix the lowest incidents.