AWS issues, affecting Heroku and others
status.heroku.com
status.heroku.com
> 10:38 AM PDT -- We are currently investigating degraded performance for a small number of EBS volumes in a single Availability Zone in the US-EAST-1 Region.
Effectively, this is exactly how it is.
> 11:11 AM PDT -- We can confirm degraded performance for a small number of EBS volumes in a single Availability Zone in the US-EAST-1 Region. Instances using affected EBS volumes will also experience degraded performance.
While it is not solely AWS' fault (cloud !== no need to do proper engineering) reddit etc. is down, I do have my doubts that "degraded performance" makes all of the ec2 instances reddit runs on go haywire and shuts down the entire website, leaving a static "come back later" html page.
Our Master MySQL server uses EBS (striped RAID across 4 EBS disks) and it is getting killed due to severely degraded EBS performance. There is definitely a major problem with EBS at this time.
Interesting to note that we issued a reboot on our Master and it went away and didn't return for over 45 minutes - we thought for sure it would have to be terminated. API calls and console access was severely restricted, so even launching in new AZ's was problematic.
They also seem to either have no issues is everything goes out all at once.
It's rather annoying as the first assignments for the Alex Aiken's excellent Compilers course are due today.
For info, Coursera seems to randomize the assignments. Specifically, in the case of multiple choice tests it seems to pick a random subset of answers to each question, and randomize the order. This is obviously a defense against dumb copying.
What it does mean though is that you can't resubmit your own work. I've re-done the assignment twice this evening (the questions change enough that you have to think a lot, even when re-doing the same assignment), only to get 500 errors both times.
Anyhow, I'm sure it'll all work out and these temporary technical problems are but minor hiccups in what is a fantastic course and a fantastic learning system. It's very, very much appreciated.
It's been about 8 years, so my memory may wrong, but I believe WebAssign used do have some sort of adaptive homework, where the questions would be selected based on your answers to previous questions (the same way the GMAT or GRE works).
I'm surprised that Coursera doesn't do something of the sort, both to avoid the problem that you're having (the 'state' of your homework would be updated each time you submit an answer to a question), and also to capitalize on some of the flexibility of digital learning that isn't possible on paper.
Not all of the courses are set up with multiple choices though. In addition to varying sets of correct responses, some answers are mathematical. For instance it is defined such that a question might ask ("What is %d plus %d", a, b), then have choices involving the quantities: ("%d", a+1), ("%d", a*b), ("%d", a+b), etc. with a different random quantity for a and b each time you start a quiz.
Interestingly, I saw one course where it had single-choice questions, with more differences than simply right or wrong. For instance one answer gave 100%, another 80%, and another 0% of the points available for that question.
I personally think that OS updates and security fixes/hardening are crucial in all cases if you're running any production site (and even for the dev environment). Granted, some of the colocation providers don't offer services like AWS VPC (and the associated SecurityGroups and ACLs) but that's pretty rare nowadays. In all cases, and Ops or a Sysadmin person is quite important for the stability of your business.
The only big advantage I see about using PaaS is the rapid and flexible provisioning of new instances in case you need them. That's why the best solution to me, so far, has been running a combination of bare-metal servers with optional provisioning of additional cloud instances if you need them. The price difference can go to orders of magnitude (even if you factor in the costs of paying somebody to be a system administrator and take care of your servers).
Edit: added the last paragraph
Traffic Spikes: Implement caching (most spikes are read-only). Implement queueing of expensive but less essential operations, so they can be delayed during traffic spikes. Make sure your CPU & RAM are 1.5 - 2x what you need. (Don't worry, it will still be cheaper than AWS)
Plain Old Downtime: Network outages happen with cloud hosts too.
Focus on development: Cloud servers are unmanaged too.
OS Updates & Security: This is no more an issue on dedicated servers than it is on cloud servers. It's an OS issue. Were you running Windows Servers. That's the problem!
Given that my sites are deployed using Multi-AZ RDS instances and yet they're still down, this takes the cake a little.
If you strip the labour cost out of the equation, hosting it yourself can be up to an order of magnitude cheaper. Of course stripping the labour cost out is only possible if you can do it yourself :)
In my case, I can (as could maciej of pinboard.in fame).
Edit: also helps if you can find a well priced data centre or two. For me, I'm planning on going with hetzner.de and ovh.net
>> Of course stripping the labour cost out is only possible
>> if you can do it yourself :)
Or assign value to your time.Admittedly, the calculation is made easier by the fact that I'm not building the next Netflix or Heroku, however, neither are 95% of the other start ups out there.
I'm fairly sure that if you have 10 instances at EC2, you're paying FAR more than $550/month.
The other benefit EC2 gives you is quick/instant deployment of even really high-end hardware if you quickly out-pace your current capacity and need another box ASAP, whereas even if you were doing dedicated servers, it still takes a hour or two for them to put your new box in a rack and get back to you with the IP and such.
Seriously though, because the costs are so much cheaper you just over provision the main machines slightly and keep a spare box or two ready to go. If that isn't enough to smooth out the spike then I guess at least you've got a nice problem to have.
More generally, other tools such as a CDN, DNS, email delivery services, Facebook authentication, Twitter feeds, OpenStreetMap. Google APIs, mapping, Analytics, and others, all rely on some third-party infrastructure, if not AWS itself. I'm not saying that all sites rely on all of these, but many use one or more.
The Web is becoming increasing interconnected in ways that increase fragility.
Talk about having all your eggs in one basket.
Many of those companies are down right now... why didn't they learn? FourSquare seems to be up - although slow. I can load Quora but didn't try to log in. Netflix, Reddit, SpringPad, seem to be down for me. http://en.wikipedia.org/wiki/Amazon_Elastic_Compute_Cloud#Is...
Of course, this could change if AWS's reliability trends downward. But it's never going to be 100%, and neither is anything you do yourself.
I've got heavily used servers at Softlayer that haven't had more than a few minutes of downtime in 4+ years, and have never been rebooted. During that same time, I ran a site out of US-East for two years (EC2 + EBS + Multi-AZ RDS + ELB), and had more downtime and spent more admin time working around significant issues with Amazon. Amazon neither saved me time nor money, nor did it provide better availability.
During that time, there was maybe network outage (at 2AM mostly) of perhaps 2 hours.
Complexity reduces reliability.
Claiming AWS' uptime approaches that of fairly standard colocation in a decent facility, is laughable.
"You are receiving this email because your Amazon CloudWatch Alarm "awsec2-WM-Live1-i-8747ece1-High-CPU-Utilization" in the US - N. Virginia region has entered the ALARM state, because "Threshold Crossed: 1 datapoint (93.13) was greater than or equal to the threshold (85.0)." at ."
Maybe their data centre was over run by mutants ?
Complete outage of an AZ shouldn't ever take services offline.
Hopefully recovering in the next hour.