Shipping is your company’s heartbeat (2013)
blog.intercom.com
blog.intercom.com
From the customer's perspective though having a fix out in a couple of minutes is the golden standard. I guess that's were team organization and planning comes into play. Having a couple of dedicated resources to address these ad hoc issues, without distracting the people working on larger projects.
Instead of bundling a lot of hypotheses about possibly-valuable changes up into a big release, you focus more on the next step in the direction you're going. You release those steps one by one. (User-visibility of larger features is controlled by feature flags, but you ship and user-test in small slices.)
In that context, the whole notion of distraction is different. You have to hold a lot less state in your head. Instead that state lives in your unit tests, your acceptance tests, your backlog, and the world at large. If a given developer is releasing a couple of times a day, it's just not a big deal to find a natural stopping point, release something small, and then pick up the next slice of some longer-term effort.
Your customers should be the highest priority, all the time. If they hit a bug that prevents them from accomplishing a task in your software, that matters 1000x more than the feature they don't even care about/use/want yet. Some big project the customer doesn't know about isn't why your customer is your customer.
Obviously trivial stuff (i.e. a misspelled word or an out-of-alignment UI element) ought not require an engineer to drop everything. But "I can't edit my document" or "I can't sign up because I have a weird email" or "I can't save or send this invoice" -- all of that SHOULD require an engineer to drop everything to fix it -- right that minute. Your customers are paying you money. If there's a fly in their soup, you ought not be saying, "If you order soup tomorrow, we'll make sure there's no fly in it. In the meantime, fuck you, I have to prepare next week's menu."
My point -- if there's a bug that's blocking a customer from doing what they're paying you to enable them to do, then that's all that matters at that moment, anything less and you're either understaffed, or your priorities are all wrong. Worry about today's customers today, worry about tomorrow's customers later.
I guess some companies have the luxury of prioritizing their roadmap over the needs of the people that actually are paying your company money -- but mine sure doesn't. We treat every customer's problem as the most important thing, because it is -- they're the ones that give us money. It would be like a restaurant ignoring a customer's request for a refill of water because they're busy planning next week's menu. You do that enough times and there'll be no need for next week's menu.
The customer isn't always right, but they should always come first.
Can't people take turns being oncall? It's disruptive to one's life to not be able to be away from a computer for more than N minutes, so only having to do this for limited periods is important.
And while people are not oncall they can work on development and finishing fixing any problems they discovered while oncall.
This sounds like a really bad idea. If it's not a critical bug impacting all your customers, do you have to fix it in 1 minute? Once your company gets bigger, you would want to do proper code review and QA, so your 'one-line fix' doesn't accidentally bring down the whole site.
Also, if you think a fix needs more care, the ability to ship quickly doesn't impede doing it slowly; but in many companies there is simply no ability to ship fast, regardless of the type of bug.
Unless it caused data corruption.
This of course is a ideal scenario that you have to work on seriously to be very close to it.
For CD to work, you can't really depend on QA for regression testing. You need to have really good automated test suites. You still do want exploratory QA, but that doesn't have to happen before release.
Code review works similarly. As your quote says, he did proper code review. It was a very small change, so it doesn't take long. Strong test coverage means code review is less about finding possible bugs and more about design issues. There's no urgency to getting those before release; people can watch changelogs just as well as they can service a review queue. And with frequent releases, developers return to important code often, meaning that design improvements happen over time.
Takes a bit of infra effort, but for a company like Intercom, that’s very easily achievable.
good luck trying this when its a backend change involving distributed threads, locks, resources, etc where production load is 100x more than staging/test.
There are plenty of backend changes where your list of complexities do not apply.
Obviously it is possible to just throw together a bunch of scripts and vaguely make it all happen, I’m not sure I’d be comfortable to do that in production.