Large-scale Amazon EC2 Outage
status.aws.amazon.com
status.aws.amazon.com
Edit: The 5 hours are what I've noticed from either personal servers or those of friends/acquaintances. YMMV.
I don't keep track of the totals but I think it you search their forums, I seem to recall someone who had monitored them with Pingdom or Wormly(or equivalent) for a fair period of time.
p.s. The last two weren't network outages - they were total power failures. Your 'node would have been hard-booted as a result.
An Amazon small instance has 1.7 GB of memory, and 160 GB of "local instance storage" (which Amazon will wipe when your instance goes down), and "1 EC2 unit" (a single core on a 1.2 GHz 2007 Xeon processor). It will cost you about $60 a month.
A similar plan on Linode will give you 1.5 GB RAM, and 60 GB storage. They will also throw in bandwidth and persistent storage, though.
Amazon tends to give you more flexibility (on demand prices, spot prices, dedicated prices, bigger instances, smaller instances, persistent storage, cheap temporary storage, and so on). But it's not as easy to use as Linode.
Linode targets people who want a cloud server. Amazon has a number of uses.
The real trouble with Amazon is the erratic performance of the EBS system which can kill an otherwise fast system. Linode's I/O does seem to be much more consistent and predictable.
Amazon's storage capacity and the cost of storage is unmatched, though. If you have a mountain of data, Linode will be prohibitively expensive.
They've essentially swiss-cheesed all our backups.
Email copy follows....
Hello,
We've discovered an error in the Amazon EBS software that cleans up unused snapshots. This has affected at least one of your snapshots in the EU-West Region.
During a recent run of this EBS software in the EU-West Region, one or more blocks in a number of EBS snapshots were incorrectly deleted. The root cause was a software error that caused the snapshot references to a subset of blocks to be missed during the reference counting process. This process compares the blocks scheduled for deletion to the blocks referenced in customer snapshots. As a result of the software error, the EBS snapshot management system in the EU-West Region incorrectly thought some of the blocks were no longer being used and deleted them. We've addressed the error in the EBS snapshot system to prevent it from recurring.
We have now disabled all of your snapshots that contain these missing blocks. You can determine which of your snapshots were affected via the AWS Management Console or the DescribeSnapshots API call. The status for any affected snapshots will be shown as "error."
We have created copies of your affected snapshots where we've replaced the missing blocks with empty blocks. You can create a new volume from these snapshot copies and run a recovery tool on it (e.g. a file system recovery tool like fsck); in some cases this may restore normal volume operation. These snapshots can be identified via the snapshot Description field which you can see on the AWS Management Console or via the DescribeSnapshots API call. The Description field contains "Recovery Snapshot snap-xxxx" where snap-xxx is the id of the affected snapshot. Alternately, if you have any older or more recent snapshots that were unaffected, you will be able to create a volume from those snapshots without error. For additional questions, you may open a case in our Support Center: https://aws.amazon.com/support/createCase
We apologize for any potential impact this might have on your applications.
Sincerely, AWS Developer Support
[i realise this is a little off-topic, but what do other app-engine users do?]
[edit: wasn't being critical, just trying to understand what others do and why]
What is the point of those status dashboards if they are not actually monitoring the health of the cloud in real time?
Google App Engine's status dashboard quickly returns to green the minute an outage goes away, hiding the overall unreliability from view.
http://reports.panopta.com/cloudharmony-borderless/server/57...
various ec2 clients affected:
- dotcloud
- dropbox
- engine yard
- foursquare
- heroku
- hootsuite
- kicksend
- netflix
- pagerduty
- twilio
Perhaps they should do the same for the website just to instill extra confidence, though.
http://www.codinghorror.com/blog/2011/04/working-with-the-ch...
There was permanent data loss on EBS volumes, which are explicitly stated to be vulnerable to data loss.
About 25% of EU west instances have been down for about 36 hours now...
Must be this.
But it really is time for Amazon to start thinking about some kind of automatic failover system.
The original us-east-1b has chronic capacity issues, meaning that it is always at or near its capacity. AWS refuses to sell instance reservations for this AZ, and it's often difficult/impossible to launch new instances in this zone.
Keep in mind that AZs are remapped for all accounts newer than a certain point in time (say, 1 yr ago). Your 1b may not be my 1b.
That your instance runs 99.9% does not say anything, you are lucky. That your instance dies in two weeks does not day anything, you are unlucky.
If you need anything more complex than delivering static content, you have to spend the effort thinking about what's important to you and how you're comfortable spending and/or compromising to meet that reliability goal. The answers for this tend to be pretty specific to your business needs and tech stack.
But in any case, I expect the PAAS people like heroku to at least start to step up their game. My apps on Google App Engine are up and I can run clojure and ruby there too.
DUH - doesn't seem to be - that was west coast, IIRC, and the only issue I see now is in Virginia.
But startups need hosts that are up 24/7. ECC doesn't give you any guarantee of uptime, and if it goes down the local (fast) disk is ephemeral. Yet, the EBS alternative which is backed up, is very slow.
The basic VPS offerings, that Linode, Rackspace, and just about everyone else offers, isn't available at all from Amazon (near as I can see.) Yet this is what startups need- local disks, a small monthly fee, and up all the time.
So, Amazon requires extra engineering--to account for nodes going down more often and to make ephemeral disks reliable or EBS performant. It also puts you on the path to lock-in, since so many of amazons services have their own unique APIs. It isn't exactly cheap, near as I can tell, when compared to, say, dedicated hosting in germany.
Its advantages exist elsewhere- if you need to spin up a bunch of machines with an API you can do this at rackspace cloud. And... that's about the only advantage I'm able to think of. Personally, I'm architecting to have some overcapacity built in and to survive a spike, because I use that excess capacity in the off hours for heavy lifting. I wouldn't try to bring up extra nodes in the morning and shut them down in the evening anyway... and I doubt that many startups are really doing that.
Possibly, I'm missing something. I tend to forget about features that aren't compelling to me, but are compelling to others. So maybe there's something that's important to these startups.
Occasional downtime just isn't that big of an issue for many startups, who are still looking for product-market fit.
Maybe you're talking about ordering servers by the hundreds, but are successful startups like foursquare doing that? Seems a few a week is what they'd add, and you could do that with most other hosts, without the added complexity of Amazon.
Dedicated hosting and rackspace cloud may be cheaper, but if you expect to be in a situation where failure becomes frequent, you'd better be on EC2.
EBS is not a unique feature, as there are multiple equivalents of EBS, along with higher level methods of replicating data, namely, couchDB, riak, cassandra, etc.
But I think you've answered my question- people use AWS because they don't realize this.
Of course the kicker is that this is in a thread about how a bug in Amazon means that the EBS volumes that were allegedly being backed up, actually weren't.
Because most startups want to gratify that vanity, they set up on EC2 because they want to prepare for the "inevitable" deluge of users without paying for an adequate pipe or server configuration around-the-clock.
Edit re your edit on dogfooding: supposedly they use the same tech, just not the same servers, so their AZs are probably separate form ec2s. Also, they're in all of their geographical locations, not just us east.
http://www.slideshare.net/AmazonWebServices/2011-aws-tour-au...
I've talked to some people about AWS about this, and the reason why they have availability zones is because they don't want to charge you the speed cost of syncing data between zones if your app doesn't need 100% uptime. Generalized replication slows down your app. AWS gives you the option of not having replication or bringing your own.
Outages are hard to avoid, but the pain can be lessened if your customers are aware of the recovery progress and you can deliver on your recovery time goals. Nothing is worse than being down, and leaving customers in the dark to start rumors that your guys are not even aware of the problem.
[1] Technically they do encrypt the files, but the keys are right next to them on the same infrastructure. Doesn't do any more good than not encrypting them.
But from a HN perspective they've received YC cash so your post is currently voted <= 0.
I like this comment.