Hot Takes on Code Freezes
jeli.io
jeli.io
Are any of these really "hot takes"?
Sadly there are many organisations out there who don't respect employee annual leave and expect their engineering team to be on-call without compensation. The usual excuses are things like "it's the holidays, it'll be quiet".
>A deploy freeze means that folks are not actively trying to make progress on a dev project that needs to be deployed. Deploy freezes need to equate to either time off, or time spent on things that do not need to be deployed when complete.
The author does a bad job explaining why. If you allow deployable work to progress without deployment during a holiday code freeze, you are just shifting, sometimes magnifying risk post-code freeze. All that undeployed work during the code freezer will increase the rate of change tremendously post-code-freeze.
There's also varying costs (not all monetary) to responding to an incident. I'm going to be very upset at my employer if I get called in on Christmas because there was a deploy going on that caused issues. First week of January, that's just expected.
1) an engineer on-call rotation for every discipline - no database engineer has to troubleshoot a broken CSS rule
2) on-call engineers get paid a big chunk of change, and even more on holidays
3) all the on-call engineers are much smarter than I am (a little harder to replicate this one, unless you're dumb like me)
I honestly kind of forget that it even exists because it runs so smoothly. Good opportunity for me to get some perspective and be grateful!
Edit: since someone asked, here are some other cool things that my engineering org does so smoothly that I forget they exist
- Very very good automated build monitoring: when you push code for review / to internal / to prod, you know what happens to it as it happens.
- Business tech is something I've literally never had to think about for more than eight seconds – laptop, idm, badge, slack, outlook, etc has worked absolutely perfectly the entire time I've been here (except, weirdly, this one meeting room that refuses to stay reserved.)
- Testing: we have a ridiculously competent QA apparatus. Our product is used by customers with extremely stringent UI requirements, and so regression testing is mission-critical, even for tiny UI bugs. Every team has a QA, but we also have a QA Architect position who helps guide org-wide decisions. Generally, QAs are also very respected and treated as peer engineers by feature developers, and I think it shows.
I work at a big SaaS company, and there are entire teams of people devoted to developer tools, so this stuff is all """basic""", I would argue; but man, it's easy to forget how much human effort has to go into making it basic, and how often we fuck up the "easy" stuff.
It's an excellent place to work, particularly as an engineer, and we're always hiring.
I find this does not match my experience at all! The number one cause of incidents are code or configuration changes. Freeze those and the incidents DO go away 95%. Of course you need to update eventually and the problems will begin again, but freezing when there is the least amount of staff to fight fires is very sensible???
of course, nothing is 100% but let's not rock the boat so that the rare traffic based incident can be dealt with by the few staff that are there.
These do reduce, but don't elimate incidents.
Unilateral change blocks, stink, but don't put support in the weird position of taking on something they can't.
The air conditioning in the server room failing when it is -24F is a possibility.
DB migrations are a different story imo, and should get more planning. I’m guessing there would be other stacks/domains where regular release intervals suit customers better though.
Why not? Thats the best part.
Companies don't want to do that. That costs money. So instead, they developed the idea of a "code freeze" which doesn't cost money (directly). In the mind of management, if there aren't any code changes, then nothing can break. Obviously, this is not the case, but that's not the point. They're not going to pay more for overtime and the marginal cost of deploying changes over the holidays does not justify the expense.
Code freezes (in my experience) don't actually reduce the number of bugs or incidents. I think in all probability they increase the number of incidents and cost money because of the warm up period.
But they do shift the timing when incidents occur, and that can be nice if your customers are sensitive to holiday sales. It's also nice to your employees. Paying your SREs a bit more for their tour of duty to the customers and employees is worth it.
I think it is ok with code freezes though.