Microsoft's Azure cloud down and out for 8 hours
theregister.co.uk
theregister.co.uk
From the article itself:
> It later added that less than 3.8 per cent of hosted services had been affected.
If this was about Google or Apple, this submit would already have been flagged, taken off the front page, and several accounts would have been hell-banned.
</endjoke, but it's really true>
It will be interesting to see if Microsoft expands on this issue and if we can learn something about it... Perhaps other datacenters are also vulnerable to cert issues in their management systems.
(or her)
Consider it a performance-based bonus.
I'm assuming that when the day rolled over, it caused an issue with an internal system at Microsoft, that had something to do with certificate dates and/or timestamps.
It's not just Microsoft that it's happening to, today I could not automatically renew any domains at Namecheap that expire tomorrow because the registry could not produce the correct expire-at date.
Leap days cause odd issues.
Year divisible by 4? => leap year; year also divisible by 100? => not a leap year unless year also divisible by 400
Seriously? Alarmist flamebait headlines about Google and Apple stuff gets posted near daily.
So here is a quiz for you, does you software know that this is a leap year? I'll speculate that someone's software didn't recognize it as such. Personally I'm always on the lookout for this sort of 'weird' thing because in the world of testing the edge cases are often poorly tested.
We have an Azure CDN backed by a Compute Instance, and zero official notice from Microsoft about this still. I've learned more about the problem from news articles than the company that provides the service. Fortunately we haven't finished migrating the rest of the site to Azure. No emails from them, nothing. Not even a tweet on their official @WindowsAzure account. Frustrating.
As someone who is actively considering building products using this platform I'm keen to see how well they manage this issue. How do they communicate during the outage, how open about what went wrong are they afterwards, do they learn from it, and so on. I particularly like the efforts Amazon have gone to in the past to show they are learning from these issues (see http://aws.amazon.com/message/65648/) I'm hopeful that Microsoft will show the same level of openness.
I'm in charge of technology for a hedge fund with 15 people and I get freaked out each time we roll out a new piece of software.
That said, I run an app on Heroku and noticed that outage - Pingdom sent me a bunch of emails, which I saw after a day spent outside. The first thing I do when I get a Pingdom failure notice is check status.heroku.com, and sure enough they had an issue and were working on fixing it. I wasn't particularly expecting there to be mention of it on HN because, first, it wasn't an enormous outage, second, it was on the weekend, and third Heroku always does such a great job of posting about issues.
I don't know if they did so, but I'd expect Intercom.io and other affected sites to post information letting people know that the site is down and why. I haven't tried it yet but Heroku allows you to specify a custom error page to be served [1], which you could update on the fly to let people know what's going on. For other outages I have let users know via Twitter and Facebook, though obviously there's no substitute for a nice up-to-date static error page.
LOL. If I had a penny for whenever the "open unix world" chose to re implement some stuff from scratch, I'd be rich. From Gnome and KDE always changing stuff (especially the KDE multimedia architecture took this to comical levels), to FreeBSD moving to Clang...
Edit: if this truly is because of 2/29, I guess anyone signing up from now on will get perfect service.
At least for the next 4 years
That said, we only rely on table storage and our instance count is mostly static.
That's typically because the 'service management service' typically kicks in when it needs to do things - allocate capacity, restart things when stuff goes down, etc. By default, it isn't touching the running apps inside VMs. There is no Windows Azure equivalent to Heroku's routing mesh to be taken down; the requests go to the VMs directly via the various networking layers.
A bit off topic: I've been trying for days to get Azure to work with .webm or .ogv files. Maybe I'm using the wrong tool (CloudBerry). I want to be able to deliver HTML5 video from the cloud but I'm without success for FF users. Luckily, my video player features a Flash fallback which is awesome.
ping www.windowsazure.com PING wamktg-prod-db-001.cloudapp.net (65.52.64.144): 56 data bytes
Request timeout for icmp_seq 0
Request timeout for icmp_seq 1
Request timeout for icmp_seq 2
Request timeout for icmp_seq 3
Request timeout for icmp_seq 4
Request timeout for icmp_seq 5
And just to add some opinion, the reason I wouldn't even consider it is Microsoft absolutely flippant ADD when it comes to online services. I have zero faith that they won't just shut it down tomorrow.
I note the other comment mentions Apple, which was a pre-release beta rumor based upon the IP that a beta iMessage lived at (since moved). It speaks volumes, I think, when such a disproven pre-release claim is still held as the example.