$1,279-per-hour, 30,000-core cluster built on Amazon EC2 cloud
arstechnica.com
arstechnica.com
The reason I ask... how hard would it be to boot up, say 10,000 micro-instances (using a stolen credit card or AWS account) to be used for a DDOS? What do you have to do before red lights start appearing in the AWS NOC?
http://aws.amazon.com/ec2/faqs/#How_many_instances_can_I_run...
If I recall, we were also limited to (10?) instances and had to ask for more and explain why.
It took a few days to get approved.
I wondered about spreading it out around the country through, seems like you would incur a lot of latency which might become a bottleneck.
So if 100% 'occupancy' on your cluster is worth a million a month, and your cost of 'owning' a cluster of this scale is a quarter million a month, then you need to do better than 25% occupancy to break even, and anything better than that and you make money. It is an interesting financial exercise if nothing else.
A lot of what we do is heavily custom, and we provide a lot of support for our customers, but we still typically beat Amazon on both price and performance. EC2 might work nicely for embarrassingly parallel workloads, but they don't have Infiniband available if you're latency-bound... :)
There are a lot more to supercomputing than just having lots of machines in a cluster. Special network connection, network topology, routing techniques, specialized CPU/GPU, specialized storage, all these can make a difference.
I'm sure you can build a special "Amazon for Supercomputing" infrastructure, rent it out to clients, and beat AWS on price/performance. Just add fast network, mix of CPU/GPU, SSD, large RAM and distributed RAM disk. Have some standard network topologies for easy configuration. Have some standard cluster layouts for different computing needs. May be having the software in place for the typical supercomputing needs. The clients just need to provide the data to a cluster and the answer will be spitted out.
However, the cash flow generated by EC2 is likely negative as I believe that (1) all the profits are re-invested in expanding, and that (2) Amazon injects money into it from other sources.
The cluster, announced publicly this week, was created for an unnamed “Top 5 Pharma” customer, and ran for about seven hours at the end of July at a peak cost of $1,279 per hour, including the fees to Amazon and Cycle Computing
Given that the article states that the entire system only ran for about 7 hours, I assume that it was one of those ideal use cases for cloud computing. So the benefit of having a disposable system adds value that is also missing in the napkin calculation. Sure, if you ran the 30k-core cluster for eternity, you might as well build your own data center. But for this case, the comparative cost analysis seems a bit pointless.
One more small nitpick: the $1279 you extrapolated from in your calculation was the peak cost, not the average cost.
"Cycle combines several technologies to ease the process" - Did that include Amazons CloudFormation services?
Porting the application to GPUs may offer a good ROI if the company intends to run this workload often enough.
One advantage of going with Amazon, is their really high-speed and voluminous ephemeral storage available per instance in addition to your EBS backed root volume.
http://blog.cyclecomputing.com/2011/09/new-cyclecloud-cluste...
In any case, I know these guys and they're not newbies. They're familiar with GPUs and what they can do. They're running the workload they need to be running.
I am not saying they are newbies for failing to exploit GPUs. There could be other reasons why they did that. For example the pharma needed the results ASAP and may not have ported that particular molecular modeling app to GPU yet. GPGPU is a nascent field after all.
I do too write GPU applications (crypto bruteforcers), just so you know ;)
http://blog.cyclecomputing.com/2011/09/new-cyclecloud-cluste...
- A single AMD HD 6990 is capable of 1275 DP GFLOPS
- A single Nvidia Tesla 20xx is capable of 515 DP GFLOPS
- A 6-core CPU at 3GHz is capable with SSE of only 72 DP GFLOPSSure, we'll never know - but it's an interesting thought experiment.