1) Did you try bigger instance types, i.e. large? I've heard at that level and above you get better networking performance, which theoretically might improve EBS performance?
2) Related to #1, I wonder if you've tried doing RAID0 on the ephemeral drives? You get two with the large instance, and four with xlarge.
3) Did you ever measure non-IO network EBS performance during the tests? I've always wondered whether using EBS heavily would slow down other network traffic to the device given there is only one interface.
4) How often have you yourself experienced EBS volume failure in your RAID volumes?
5) When that happens, what happens to your volumes and instance? That is, what do you use to monitor when the RAID volume degrades? Does it usually take down the instance immediately or only after some time? If it just becomes really slow does that throw alarms or does the application just become really slow?
6) Finally, what is your current procedure for dealing with a volume failure?
Well that ended up being a lot of questions. I'd really appreciate any answers/insights your or anyone else could shed on these questions. I've been reading all these posts but it feels a bit like reading tea leaves.