Maintenance windows are a mistake
infoworld.com
infoworld.com
There are many ways to approach a problem, and there are good and bad ones, but often you have to make a trade-off. I would appreciate a more differentiated view.
The author makes the math on "high cost of maintenance windows", but leaves out the expenditure side to it. And what is the hand-rule for adding a nine behind the availability? I vaguely remember exponential cost increase. "No downtime, ever", requires next to infinite cost.
The author suggest a false dichotomy. You can have maintenance windows and while aiming for zero downtime. The maintenance window serves the customer a higher predictability on when to expect failure. No failures are obviously preferable, but it is never the question if to cause failures or not, but where you spend your resources on.
This might make sense for your inner city loop, with trains passing by every 30 seconds (and it's gonna help when one train inevitably breaks down). However, it's a totally unreasonable cost for a station deep in the country where trains only stop 4 times a day.
To extend your analogy, it might be a terrible idea to start doing maintenance on a single line track that sees high usage in a holiday season just before those holidays where everyone wants to get away or get home as do your staff. So you might implement what is usually called a change freeze.
On the other hand you might have a different time where demand is low. Shutting down the line if it comes to that will annoy some people but not as many. So you plan your maintenance then and have people ready in case things go wrong.
Software is cheap to produce and quick to change. Building a "third track" on the fly, as needed, is exactly the kind of thing we can do that actual engineers can't. They probably wish they had the ability to do things like that.
For some apps, the big issue is when you schedule the downtime to happen. If you are truly customer-centric, you should schedule for a time that causes the least disruption for your users. Of course, a decision like this depends on other factors - the number of users likely to be affected, how critical the app changes are, costs, etc.
I've noticed a number of companies in my country taking sites and apps offline in the middle of the workday when they're very likely to be in use. To me, this is unacceptable - the work on the app should be done after hours.
I specifically took a job that didn't require after-hours work (I have a family which I prioritise) and we did upgrades in working hours. Once in a blue moon a site would be down for a minute or two, but the trade-off was obviously worth it for my public sector employer which generally struggled with recruiting engineers.
This is spot-on. I see this attitude often when there is a lack of data being exposed for observation and it "just feels" like ____.
I occasionally get notices of planned downtime from services I use. It's almost never an issue for me - the outage windows are usually relatively brief, infrequent, and off-hours. Not every service has the luxury of even having "off-hours", but simply making planned downtime infrequent goes a long ways.
I don't get this - who wants 'predictable failure'?
If you're a B2C business, this is completely unacceptable - consumers don't plan their google searches around your preannounced downtime windows.
If you're a B2B SaaS business, well... your customers downstream have their own downtime to manage. If they have two vendors, each of which have 'scheduled maintenance windows' that don't overlap, then their combined availability just dropped dramatically.
Far better to build systems that are generally resilient to being down - queues for offline processing, idempotent operations, retriability... you'll wind up needing them to handle scheduled downtime anyway, and once you have them, you can switch to 'as needed' maintenance with no loss of service.
Also, some changes are one off and significant (like some database changes), they will hopefully go smoothly, but it is hard to fully prepare for every possible issue.
People need to voluntarily be disconnected from time to time or experience it due to scheduled down time. Maybe follow some international (or local) holiday schedule.
Life will go on and people will find solutions to the time when things are off.
You should not have to rely on Netflix for entertainment or Facebook for communication, google for information, etc.
People will be better prepared for emergencies this way and won't be left as helpless.
In a smaller IT shop (~8 people) maintaining all services for a 500 person 9 to 5 business maintenance windows provide a structured way to to perform all required tasks without the additional overhead of high availability, especially in the cases when it is not necessary.
The caveat is excellent communication regarding the plannee maintenance (what, when, why and the anticipated impact) and when systems will be available again.
But I am not a "recognized thought leader in cloud computing" as the author is, I just know what has been practical in my experience.
You and your customer can save a lot of money by doing so, which is important especially when the customer does not post billions in profit every year or swims in VC money.
Maintenance windows are out of working hours most of the time, so it's not hard to not impact business.
Isn’t that obvious? Did it really need to be said?
What the article misses is that in most cases where I worked with clients, downtime was not necessarily for the customer mainly, but for the internal non tech team. You need planned downtime to: have a proper recovery plan in case whatever you’re doing doesn’t work out and you gotta revert, or so that you can work comfortably. I had downtime at different moments day and night and planned downtime during the day or night; planned always wins on so many levels.
If you think planned downtime is not downtime, it’s not the strategy’s fault, it’s your own thinking.
The whole thermostat thing is really not worthy of addressing. Your technical dependency justifies a feeling, but an opinion is a bit much. At best you can call it a half baked thought. Definitely not worth an “opinion piece”
It really comes down to the cost of implementing "Super HA" for everything. If that cost is more than the tangible loss of having 30 minutes down each month, it's not worth it.
And it's not just IT stuff that has to be "Super HA", it's everything. Updates need to be atomic. Every database change needs to be able to be half-done so staged rollouts can take place on application servers. That requires 10 to 100 times more effort, and therefore cost. Not to mention testing, certification, etc. It all comes with baggage.
I run a bunch of very small sites on a single Digital Ocean droplet. It has no HA whatsover, it's just a small VM running nginx. Yet, somehow, it has better uptime than every single one of the 10 biggest sites last year (30 min downtime total for 2020). Sure, I could beat that down to zero by using a load balancer and multiple servers, replication, shared storage, all the shiny stuff, but it's just as likely that something could go wrong because of the added complexity, and take me down for an hour or more.
There's never a single Perfect Solution to every problem. That's why we don't build suspension bridges across creeks.
Instructions how to
* upgrade server-applications quickly without user interruption
* to create autonomic applications on client-side
* how to change data models in the background
Desktop mail-client like Evolution or "git --everything-is-local" are good examples. They work autonomic and can sync with servers anytime. The examples of the author are awkward exceptions. Real world ones are available. For example Garmin Connect or Google Mail. Garmin Connect cannot do anything autonomic and we've seen an outage last year over a complete week. The same for Google Mail, it cannot download all mails in an mailbox. On a mobile device with 128 GB disk space. You see the irony? A mobile device which will suffer concnetion loss often which a big disk? Leets look at K-9 mail or Evolution: Download all? YES! Sync when possible? YES!
That is my recommendation. Make your local application autnomic and use servers to sync when possible. Of course that cost time in development. And it should - reliability costs. Making a server un-failable isn't possible.
By the way, I'm interested in quick and better migration best practices.
Maintenance windows suck, but they work, and are predictable. It's easier to upgrade a database by bouncing it with a bit of predictable downtime, much much harder and expensive to do it live with zero downtime.
Edit: not to mention that humans understand and are usually incredibly forgiving. But they are rarely forgiving when surprised with outages.
Although in the company I used to work for, I can remember there were stories of people who liked to carry out some customer journies at 3am in the morning.
The only environment other than maybe mainframes that supports hot reloading is Java and that has limitations.
Linux has zero problems with you overwriting a binary (by unlinking and replacing it) while it's in use. Sure, the old version will keep running, but it doesn't prevent the update, and you'll get the new version when you restart it.
Even then there’s still downtime when you restart the executable (could be seconds, but also minutes). Preventing that is possible, too (even on a single machine), but the cost often doesn’t make it worth it.
And nitpick: I think you know it, but for others: “unlinking and replacing” is essential. “Overwriting” (keeping the inode number the same) can be problematic.
Yes, that is why /usr/bin/install exists.
No, in Android you can keep using the app right until the actual swap happens. Then the app restarts (and presumably a symlink is put in place in between). In practical terms this is zero downtime.
That the case with a mainstream operating system. Only one. All the others can count uptime in years even with regular application updates.
Every major Debian/Ubuntu upgrade asks me for restarting services due to a libc upgrade and warns of potential instability as a result of not doing a restart. Not to mention upgrading systemd, which tends to require a reboot.
I'm pretty sure if I had smart stuff at home, I'd be tempted to check it all the time. Better avoid that.
Maintenance window is the morning during coffee, before my wife wakes up :D