Cloud Server Reboots
status.rackspace.com
status.rackspace.com
I was working as a developer at Unisys (or it might have still been Burroughs), down the hall from the OS developers (for one of Burroughs' 3 distinct mainframe product lines, each with its own cpu architecture, operating system, system software, etc.). At that time, our mainframes could be patched while running, without re-booting; re-booting was considerably slower than patching. As a result, as each new OS patch was developed, it was applied on the fly to the machines in our dev environment, and they hadn't been rebooted for something like six months. It turned out that it was possible to have a sequence of patches, which when applied on the fly to a running system caused no problems, except that when it came time to reboot, it would crash. You can imagine the fun of trying to debug which change (or combination of several changes), out of many months of development work, had broken the boot process! These dev machines were shared with other teams (compiler developers, database engine developers, etc.), so the productivity hit waiting for our dev machines to be available again was substantial. After that, the policy was that the OS developers took turns coming in early one morning a week to reboot the dev machines, so in the worst case they'd only have to track down which of the last week's worth of patches might have broken the boot process.
Not too long ago I had a 36-hour straight marathon sysadmin session when a routine update happened to break my particular stack without advance notice on a web server used by a bunch of customers. I think my eye still twitches when I think of it.
Usually the next thing people say is, "you should set up x, y, and z tools for better redundancy / failover / load distribution / backups / sysadmin management / etc." Again, they're not wrong, but none of those things directly makes you more money, and making time for them can be difficult, especially since so few of them can be set up as easily as they claim on the box.
It sure would have helped those guys working at Chernobyl.
pay particular attention to this reboot if you recently
did "apt-get update" or similar to take care of the bash
problem
I don't understand. apt-get update won't install new software on reboot. That just updates the index in apt. There's no other way to update bash without "apt-get update && apt-get install bash" and that won't impact anything on reboot.Your updated mysql config file sounds like you apt-get upgrade'd, installed a new version of mysql, or something like that and the changes didn't go into effect until the mysql process was restarted. I got bit in the ass with that same thing when I apt-get upgrade'd.
cries softly to self
What is it to fix then?
They suggest taking backups/snapshots of your instances before the reboot window. Given the throughput required to push multi-hundred GB images from public cloud servers to Cloud Files for storage, I am willing to bet that the backup network is maxed out and will stay maxed out until the outage.
I wonder if Rackspace found out about the rumored Xen exploit from the same people that told Amazon, or if Amazon told Rackspace but waited a little bit to make it more painful for Rackspace's customers...
The next question then is, who does have access to this? It definitely needs to be kept wrapped-up at least until the large vendors are patched, but who gets to decide who's a large vendor?
http://www.xenproject.org/security-policy.html
They also spell out the policy for who is eligible.
Xen has previously had problems with people leaking embargoed security issues http://www.gossamer-threads.com/lists/xen/devel/248239
Even with our company which has in in-house ops team and our own colocated kit, we expect process failures and restarts and plan accordingly. No sleep is lost even when something explodes at 2AM (other than for the ops team, who in this case you contracted out).
I want to be woken up if something doesn't come back, not if it does or even if it has gone into limp mode.
Seriously you're looking at the wrong end of the problem.
But others are not like you. Systems are not always 100% foolproof, people don't have comprehensive DR plans.
Is it that hard for you to understand why someone may want to be up when there is a reboot? Like seriously?
Someone who thinks they can test all possible failure scenarios is severely lacking in imagination.
Sure there are more possibilities, but these should all be covered with a complete fallback DR strategy i.e. if not identifiable cause X, Y or Z then assume the worse, snapshot everything for debugging later and fallback to a complete restore from scratch.
The trick is to make X, Y and Z self-healing and part of your architecture provisioning.
I get the feeling most people here haven't dealt with large deployments and multiple site redundancy, five-nines reliability requirements and extensive DR planning, because the above is pretty obvious to those of us who have.
At some point on Sunday, I'm going to be picking the pieces of our entire stack.
Rackspace doesn't even offer anything like availability zones.
The last major maintenance they scheduled was over the July 4th weekend -- wasn't happy with that one either.
Turns out it's a moot point anyway, because cells are client-visible, and have been contributed upstream to OpenStack as a Nova extension.
The only way we find out about cell locality is when our account rep gives us an updated spreadsheet. We've inquired about this several times.
If you happen to actually use the Rackspace public cloud and know a specific API call we're missing, I'd love to hear it.
First of all you communicate all details so I know how I will be affected. Secondly you don't shut down my service ever, for any reason, other than lack of payment.
If you can't do those things you don't get to claim to have excellent support.