We use EBS extensively on our infrastructure. Occasionally an EBS volume will fail. We ran into an issue where a volume had a spike in IO load as I think was the case here.
EBS volumes are not magic, they are just chunks of physical disks. Disks fail. You should have the system architected so that you can handle such failures.
Here are a few things you can do (we don't do all of these, but enough to assure we won't lose data or have downtime in the case of a failure.)
- Mount several EBS volumes, use raid.
- If its a database, set up a failover node on a separate EBS volume.
- Take regular snapshot backups.
- Take regular full backups to S3.
It's also very important to have everything highly automated. If an EBS volume fails for us, its one command to switch to the failover node. If that doesn't work, its another command to spin up a new machine off the last available snapshot with a few hours of data loss (worst case scenario.) Everything is highly monitored with nagios + ganglia so we know when bad stuff happens.
The two or three times we've had issues with EBS we were able to either switch to a failover node or take a snapshot and mount a new volume from there. I haven't set up RAID on EC2, but I'd imagine this also a very good route to protect your data.
Remember, the cloud isn't magic. The only advantage you get with the cloud is rapid provisioning and unlimited capacity if you need it. You still have to build a shared nothing, reliable architecture within the framework the cloud gives you. We've found EC2 and EBS to work out very well, but of course there were growing pains as you learn very quickly where your single points of failure are! I get the sense that the overall reliability of resources such as instances or volumes, on the whole, is definitely lower than what you'd expect in a standard hosting provider, whatever the reason may be.
Edit: Of course, you could also load your data on a distributed data store like Cassandra as well that handles some of this failover and replication magic automatically, too.