EC2 I/O
blog.scalyr.com
blog.scalyr.com
Here are links to a couple summaries of the analysis I’ve done on EC2, Rackspace and HP. I plan on writing a blog post regarding this analysis soon.
Disk Performance: The value columns is a percentage relative to a baremetal baseline 15k SAS drive, where 100% signifies comparable performance. Benchmarks included in this measurement are fio (4k random read/write/rw; 1m sequential read/write/rw), fio – Intel IOMeter pattern, CompileBench, Postmark, TioBench and AIO stress:
http://dl.dropbox.com/u/20765204/1012-disk-io-analysis/disk-...
Disk IO Consistency: The value column is a percentage relative to the same baseline. A value less than 100 represents better IO consistency than the baseline. The value is calculated by running multiple tests on an instance, measuring the standard deviation of IOPS between tests, and comparing those standard deviations to the baseline. Testing was conducted over a period of a month on multiple instances in different AZs.
http://dl.dropbox.com/u/20765204/1012-disk-io-analysis/disk-...
The same issue applies to network performance. I've seen very expensive EC2 instances that couldn't even push 50 Mbit/s to the net, while instances of the same type could at least do a few hundred Mbit/s. AWS' answer was always to simply buy even more expensive instances, so less people are sharing, but that's a terribly costly answer.
I'm doing bonded (802.3ad) 2x1 Gbit/s connections on all servers, because that's what I wish EC2 had.
Multiple customers, with highly varied workloads, sharing the same physical server hardware is simply a fundamentally flawed idea. IMHO, it only makes sense to use a VPS for very small personal projects, where you don't want to justify ~$140/mo in server costs.
EC2 was a really novel thing and it brought lots of great technology to the scene, but they made a few fundamentally wrong choices.
The graphs are very well done. That amount of data would have been incomprehensible if not for your carefully thought out charts.
I'll see whether we can work larger RAID configurations into a follow-up.
We'll very likely do a followup post to test provisioned IOPs. If anyone has other suggestions to include in the followup work, please let us know.
$/iops between storage types is the subject of the "Cost Effectiveness" charts near the top of the post, but I suspect I'm not quite catching your meaning. What do you mean by $/iops/storage?
One question: on the throughput graphs, I understand why you normalized them per graph, but were there any differences between graphs (particularly in terms of EBS vs. ephemeral) that would be sufficient to drown out the variability within the throughput graphs?
If you only provision 1TB (or larger if they have them available now) EBS volumes then you'd have spindles dedicated to you whereas with smaller ones there might be a lot more variations because you share.
More background: http://perfcap.blogspot.com/2011/03/understanding-and-using-...
A 4k write has to be synced to all the disks unless they have a <=4k stripe size AND are using RAID-0 AND are using stripe-aligned IO ops. It's also possible they use 4k writes to cache that end up forming large dirty blocks which the OS then syncs as larger I/Os. But that would be measuring something else than the benchmark claims.
Nothing to see here, move on...