18 hours, $33K, and 156,314 cores: Amazon cloud HPC hits a “petaflop”
arstechnica.com
arstechnica.com
I wonder if people miss this part. For easily parallel projects, Amazon is great since any core can make progress regardless of the other cores state or reachability. But if there is any inter-dependence between the cores (I use the term engtangling but others have called it co-dependence) the number becomes limited by the bisection bandwidth between cores. Poorly scheduled (like randomly scheduling on a core) is the worst case, if you can keep the bandwidth between entangled cores high, even if the system bandwidth is low, you can improve things (dramatically some times). But at the end of the day its the level of sharing between cores that will tell you if you can do your problem in the 'cloud' or if you are going to need a local data center.
Anyway, the real problem from what I've read in the literature is in problems that require network IO, but going back now, most of that seems to predate Amazon's HPC cloud service and even a year ago, some people have found pretty good scaling there: http://www.computer.org/portal/web/computingnow/content?g=53....
they count AWS "vCores" which is 1 HT ( half-core).
Does anyone know what problem/model was used for the calculations? I find it strange that there are many details about the hardware, but no information about the problem set in the field of organic solar cells.