Limiting breakage with a software deployment checklist
blog.gojekengineering.com
blog.gojekengineering.com
Also this?
> And once you fail a build, then every team member in your team has to do deployments and go through the deployment checklist.
Sounds like there's a piece of context missing from the section before it. You have to do the checklist to deploy, and if you fail once, then every member of your team has to do the checklist as well?
Then the ticket is picked up on merge to master which creates a change list. Everything gets deployed to QA where it's tested. QAs go over the changes and approve it, then developers push to production. Afterwards a random developer from a different team is picked to do a post-deploy review which spreads knowledge and sometimes picks up other changes.
If something can be automated it should be :)
So they have no change management process in place and are basically hacking on it 'till it works. Very professional.
Does not look like they ever heard of the capability maturity model, either.
I'm very much inclined to defend their methodology and position on this since that is their core area of expertise, and I've experienced it working on a very large scale (tens of thousands of servers) rather than a bunch of hacking efforts at some company on the InterNet.
https://www.amazon.com/Capability-Maturity-Model-Guidelines-...
https://www.amazon.com/Managing-Software-Process-Watts-Humph...
Peer review (as in git workflows with pull/merge requests and audit), automated testing, pipelines, CI & CD are all just a modern form of change control, where we've remove the checklists from humans.
Plus, if they provide a good service, people will forgive them. Twitter was terrible in the first years it was growing in that regard, yet people kept coming back.
https://sendgrid.com/blog/change-management-keep-it-simple-s...
Doing it formally for every deployment seems like it would kill productivity.
For the final push into production, the developers have to be online (2:00 am Eastern), along with QA and OPs. QA or the developers can abort the deployment [1] for any reason, and rolling back is trivial. So far in my seven years at The Company, I've had to abort a production deployment once (yes, I noticed an issue and aborted the deployment---it was totally my call).
[1] Our customers are the various Monopolistic Phone Companies. We have scary SLAs. We get approval for deployment from them. Downtime costs us Real Money. I don't get to deploy stuff all that often (stuff that doesn't talk directly with the customers is easier to deploy---unfortunately, most of what I work on talks directly with the customers).