Evaluating Amazon’s EC2 Micro Instances at DocumentCloud
blog.documentcloud.org
blog.documentcloud.org
This is especially important in AWS, as the dirty secret of small instances is that they occasionally vanish without warning -- in our case this has happened 4 times in 18 months.
Is having your OS on an EBS (as is required for micro) a dangerous thing, because the instance vanishing without warning might result in a corrupted EBS drive, making restoration more complicated than spinning up a new instance?
When the Micro went to OCR 51 pages of document text, it spent only 6 minutes of actual CPU time to do it. Unfortunately, those 6 minutes of CPU time were spread over a real wait of 52 minutes, due to either other users on the machine, or Amazon caps. That's not so great for an intensive burst.
Perhaps it would work if the CPU burst lasted for seconds instead of minutes, but I wouldn't bet on it. It would be a nice benchmark to try...
In a Rails app (for instance), the overall performance of an application is often constrained by how many instances you are able to run in parallel. If a 32-bit box can run 50% (or 100%) more instances of your application than the equivalent 64-bit box with the same RAM, that's great.
My own best guess is that it's that classic bugaboo: Historical reasons. They've got a big installed base of hardware, software, and ops procedures to support 32-bit small and medium instances, and thus the marginal cost of continuing to run that hardware is lower. Given the choice between (a) a big program to compel all their existing 32-bit customers to migrate to 64-bit infrastructure, and then writing off all the racks full of 32-bit equipment; or (b) maintaining and building out the datacenters and infrastructure that they've already got established, they tend to choose option (b).
Of course, given that the price of compute hardware continues its inexorable fall, in three years Amazon will offer 64-bit at an intermediate price point: The expensive 64-bit large instance of today will be the cheap 64-bit large instance of tomorrow, and people who want real power will be paying for 32xlarge instances.
With EC2, I can have a server half the price of Rackspace's cheapest instance, but with theoretically unlimited storage space. I could use that e.g. for backing up server data, or as an NFS share.
The one major limitation of Rackspace's service is disk space. It's severely limited, and there's not much you can do about it.
You get /dev/sd[b-p][1-15] as mountable EBS volumes. Thats (15 * 15) * 1TB, or about 225TB. You may even be able to go higher than "p", but I haven't tried. I'm also not sure, you may even be able to attach somewhere other than /dev/sd* .
Either way, after that, your EBS bill is $22,500/mo, just for storage, so you might be able to figure something else out for that cost.
But you CAN attach that to a $14/mo EC2 instance, which was my original point.
Edited to add:
Note that the default limit is 20 EBS volumes per EC2 account. You can request an increase from Amazon if you need more. http://groups.google.com/group/ec2ubuntu/web/raid-0-on-ec2-e...
Still, 20TB is "practically unlimited", for most applications.
Also, in response to your question: I can confirm that EC2 complains if you try going above sdp, or anywhere else.
If you have a system with 512MiB ram on a box and a system with 2048Mib ram on the same box, and they each have 1 vcpu, and they are weighted appropriately, when the cpu is idle, the 512 system will have just as much cpu as the 2048MiB system. When there is contention, the weighting will step in and give the 2048MiB system more CPU, though.
(of course, I believe that amazon does not mix instance sizes. Even so, assuming an even ram to cpu ratio, which is a pretty good assumption, the 2048MiB ram system on a system full of 2048MiB guests will have much better worst case performance than a 512MiB guest on a system with a similar ram/cpu ratio, even though if they both have 1 vcpu their best case performance will be about the same, assuming the cpu cores on each hardware are of similar speed.)
Really, if you are going for imperial data, you should do it the way Eivind Uggedal did it. He repeatedly ran a benchmark and recorded results over a very long period of time. Doing it just once is going to get you a pretty random result, and on average, is going to make a small instance look much better (cpu wise) than it is in the worst case, and really, you usually have to design for something fairly close to the worst case.
sudo gem install docsplit
Then just `time` the runtime and compare.-bash: gem: command not found
The article made it sound like they only ran it once on each instance, which means there are many factors that could have affected the results such as disk caching, other processes, or high load.
As one of my physics professors uses to say (paraphrased): A measurement is meaningless without an estimate of the error (+/- 0 meaning).