Mailgun downtime resulting from Rackspace cloud server reboots
status.mailgun.com
status.mailgun.com
Also, since Rackspace didn't communicate it that effectively: if you're on their first-gen infrastructure (your servers have an asterix next to them in the list), you won't get rebooted. I had to hear this from a fellow customer who heard it from support, as I was searching my inbox frantically wondering why they hadn't given me a heads-up about an incoming-any-minute-now reboot yet.
Unless I missed something here, I didn't receive the notice until going home for the weekend that this was happening this weekend. Hopefully someone from Rackspace is reading this: If you're going to reboot every VM on a weekend, tell us before the weekend, please.
Rackspace PR: lets release the notice Friday night so we don't get bad press.
Although I do agree that more notice would have been helpful.
The engineering team probably spent some time running tests and scribbling on whiteboards, trying to prove that the boat wasn't going to sink. In hindsight, they should have just sounded the klaxon and started handing out life jackets, but you know what they say about hindsight. And there are lots of reasons why the typical engineering organization struggles to accept the inevitable and call for an evacuation. Nobody likes Cassandra. Everybody wants to be a hero. Didn't you say this boat was unsinkable? It's hard to get all the decision-makers into one room. The show must go on. It isn't obvious that this complicated problem leads to our certain doom. Et cetera.
The key to making these things go smoothly is the Chaos Monkey, a.k.a. "conduct constant drills of your emergency responses". If you don't rehearse the response, you shy away from trying it. AWS halts or reboots EC2 instances all the time, and lo and behold, when it comes time to reboot all EC2 instances they don't flinch. Or they flinch less visibly, anyway.
While their status page was somewhat helpful, I find it absolutely absurd that they can only update it once every 60 minutes to cue customers in. In addition, their Cloud Control Panel doesn't reflect the reboots. When a VPS goes down for a reboot... the CP shows the server as "Active - Online and functioning as properly". Thankfully great third party monitoring services (using Scout) exist so they can notify in-place of Rackspace's incompetency.
As someone who shells out a significant amount of money to them each month, this is pretty disheartening. That being said... I seem to have survived the Great Rackspace Reboot of 2014 and can only hope they handle the next event better.
Our Heroku MG logs indicate that all messages are getting "Delivered", but that doesn't match reality. We've been testing with our own accounts -- ones that receive copies of all automated emails, as well as our personal accounts.
The Heroku MG logs says "Delivered" for all the emails.... but using 4 different addresses across 4 different carriers confirms: a VAST number of emails (since noon today) were not actually delivered. The only change in configuration: MG's downtime. I seriously hope all of these "Delivered" emails are re-sent. If someone from MG could weigh in, that would be fantastic! (We have a ticket filed, but email also in profile).
At work every machine reboots at least every month. Everything is designed to cope with that reality.
Count your blessings. Not all of us work with software so well designed. For us, at least, a server going down can be anywhere from an "eh" to a desperate need for attention, depending on the exact VM. Some technologies just don't take it that well. (Including many that are on HN.)
Further, every server went down in a very tight window. If you weren't replicated across regions, for all you knew, every VM might be rebooting at the same time. I know you should be replicated across regions, but again, count your blessings.
And not every organization has the budget for multiple machines, many run on just a single machine. Yes, that means accepting downtime may happen - but that doesn't mean being happy about it.
I suppose it takes a fanatic to justify working week scheduled down time.
Anyone recommend a good dedicated server Rackspace competitor?
What poor execution on this...Xen updates can't be that rare that this is their first rollout.
I just recently migrated a 180 GB instance from one KVM compute node to another in 3 minutes.
From memory[1], you had to be migrating between the same architecture, and if they're anything like the smaller providers they're running less bigger nodes, and thus it's actually a lot quicker to just push the power button.
[1] I worked with Xen years ago, my information is no doubt out of date.