A few high end cloud server IO performance comparisons
cloudharmony.com
cloudharmony.com
We read so many times about very unreliable performance. Sometimes it's ok and sometimes it's really, really bad.
Without any kind of time period to continuously run a benchmark in, this doesn't really help. For all we know, the first placed service was just having a good day and the last placed a very bad one.
As a hobby, I've been running long-term benchmarks on a handful of cloud services, for exactly the reason you suggest -- performance varies quite a bit over time. You can browse through some of the data at (http://amistrongeryet.com/dashboard.jsp). The UI on this site is abysmal, and it only shows 30 days of data (I've actually collected almost two years), but it still gives some flavor of how much variability there is. For instance, check out this graph of SimpleDB reads: (http://amistrongeryet.com/op_detail.jsp?op=simpledb_readInco...). I've blogged on the data from time to time; for instance, (http://amistrongeryet.blogspot.com/2010/04/three-latency-ano...).
If there's interest, I'll work on making this data more accessible, adding documentation, and providing access to the full two years. In any case it's limited by the fact that I'm only probing AWS and Google App Engine, and only one or two instances of each. What I'd really like to do is open this up to crowdsourcing -- as I discussed a while back at (http://amistrongeryet.blogspot.com/2011/07/cloudsat.html). If anyone is interested in participating in a project like this, let me know!
I spun up a storm-ssd-3gb instance and ran bonnie++ on it. Results here: https://gist.github.com/2069845
If I'm reading right, its ~868MB/s seq write (CPU bound) and ~594MB/s seq read (CPU bound). Not sure how to read the random IO results. The bigger instances would probably be faster still (more CPU).
So the Storm on Demand SSD instances seem to be blazingly fast.
Why are the bars in the chart in a different order than the lines in the table? Maybe have the different amazon/xyz-provider be different shades of the same color?
P.S. I expect someone will ask for more specifics, so here are a few. First, Bonnie++ sucks. Many of the numbers it produces measure the memory system or the libc implementation more than the actual I/O system. I've seriously gotten more testing value from building it than running it, so its very presence taints the result. Second, fio/hdparm/iozone might be redundant, according to which arguments are used. Or the results might be non-comparable. Either way, the aggregate result could only apply to an application with exactly that (unspecified) balance of read vs. write, sequential vs. random, sync vs. async, file counts and sizes, etc. Did they even run tests over enough data to get past caching effects? That's particularly important since they used different memory sizes on different providers. Similarly, what thread counts did they use on these different numbers of CPUs/cores? Same across all, or best-performing for each component benchmark? With such sloppy methodology anything less than an order-of-magnitude difference doesn't even tell you which platforms to test yourself.