Jedberg on Reddit's most recent outage
reddit.com
reddit.com
I wrote a long post about it here: http://www.deserettechnology.com/journal/reddit-the-open-sou... .
[1] http://www.reddit.com/r/blog/comments/g66f0/why_reddit_was_d...
There are co-lo provides that have data centers in different places.
In practice and on cheap hardware, networked storage is flakey and has umpteen failure modes. Database replication is even worse. Both require babysitting by developers or sysadmins and hours to repair when it goes wrong. What is the point outsourcing hardware and scaling to EC2 if you end up with even more work to monitor and keep fixing the infrastructure you build on top?
Looking at their past outage response:
http://blog.reddit.com/2010/01/why-did-we-take-reddit-down-f...
Money quote: "In response, we started upgrading some of our databases to use a software RAID of EBS disks, which gives drastically increased performance (at a higher cost of course)."
RAIDing EBS disks seems like a really really BAD idea. There is a non-trivial failure rate of any single EBS disk, and if you RAID them together, your failure rate of the RAID will consequently increase. Am I understanding that correctly?
If they fix that, could that be a 'silver bullet' to fix these outages?
It seems like that would make a lot more sense for Reddit, since I/O is so slow and flakey on EC2 (from what I hear), and it's not like they're really taking full advantage of the elasticity that EC2 provides (by massively scaling up and down to fit major load variance)
Since local disk is temporary and can be lost any second, don't they still have to use EBS for persistence?
EBS is a pretty bad product regarding reliability and consistency of performance unfortunately, it's better to design your system without it.
If you can't use any persistent storage, then these machines become pure processing nodes. Which in this case, I feel like it would be better to not design your system with EC2 at all. :-(
To me having EBS made EC2 a very powerful solution compared with their competitors. Not having durable and consistent EBS otherwise makes no point in differentiation and serves purely as non-functional fluff we end up paying extra for.
My general take on this issue is that if you're running your app on EC2 and your persistence medium is something that's also on EC2, you really have no ideal high availability scenario. Of course, even in my case, if SimpleDB and S3 go down, I'm still in trouble, but at least I have the option of throwing Akamai in front of it.
I don't see how this helps?
i believe the point is that instance storage is more reliable than ebs.
EBS is really just like any other NAS/iScsi vol and it wouldn't be such a problem if they did what they're supposed to. That is, be consistent in read, write, and durability.
You might think your safe with reserved instances too, thinking you've reserved the dedicated time with EC2. Well what happens when the entire network stack goes down or block storage the reserved instance rack depends on goes down? So does all 10 or 20 reserved instances you had up, at once.
Also, it was this very replication/snapshot mirroring feature of EBS that cascaded into (network,etc) congestion.
Check http://joyeur.com/2011/04/22/on-cascading-failures-and-amazo...