Nobody ever got fired for buying a cluster
research.microsoft.com
research.microsoft.com
That's 168 GB RAM total, within a single upgrade unit of their 192 GB single server, suggesting the problem is dominated by RAM.
If I had $6638 + 2640 = $9278 to spend on computing hardware from NewEgg, how about:
1 at $1000 of HP ProLiant DL360e Gen8 Rack Server System Intel Xeon E5-2403 1.8GHz 4C/4T 4GB http://www.newegg.com/Product/Product.aspx?Item=N82E16859107943 http://h10010.www1.hp.com/wwpc/us/en/sm/WF06a/15351-15351-3328412-241644-241475-5249570.html?dnr=1 (12 DIMM slots)
4 at $70 of Kingston 8GB 240-Pin DDR3 SDRAM ECC Registered DDR3 1333 Server Memory Model KVR13LR9S4/8 http://www.newegg.com/Product/Product.aspx?Item=N82E16820239540
1 at $54 of Seagate Barracuda ST250DM000 250GB 7200 RPM 16MB Cache SATA 6.0Gb/s 3.5" Internal Hard Drive http://www.newegg.com/Product/Product.aspx?Item=N82E16822148765
$1334 ea server, 7 servers = $9338
So we could get 7 of these low-end name-brand 16 GB servers for the same money to give us 224 GB RAM.> MR++ runs on 27 servers whereas the standalone configurations are a single server running a single-threaded implementation
Sure, nothing will beat a single system at message-passing algorithms when the entire graph fits in main memory. But when the dataset outgrows that (and it will), we can triple the RAM in the empty slots, or add more servers in units of $1334 instead of having to rewrite your whole analysis.
I'm confused. I think the link changed to a different version of the paper. The numbers I went by aren't even in there any more!
I think the paper has a good point: scale-up should be definitely considered. I see systems that support 1.5 TB of RAM now... do you really expect your dataset to grow beyond that?
And as long as the dataset fits within 1.5 TB of RAM, scale up will be more scalable than scale out.
> I see systems that support 1.5 TB of RAM now... do you really expect your dataset to grow beyond that?
According to Figure 1, about 20% of their jobs are already beyond that. If you spend your entire hardware budget on a single server and all your analytic processing jobs end up being coded with the (gasp) single threaded assumption, what are you going to do next?
Most of us have been there before because it's where you end up by accident ("Do you think if we max out the RAM on the database server it will return to acceptable performance?") This is why people love parallel map-reduce based systems.
I'm more concerned about another problem. There are people out there who think that they can scale out with Hadoop cluster built on top of Atoms (or more recently... ARMs). And while yes... there are some applications where that is a decent scaling strategy, I doubt it works in most cases.
But yeah, I know that wasn't your argument. I'm just musing a bit on random stuff.
But at the moment, the evidence is solidly in favor of the typical 2U Xeon. (by nature of ARM servers barely exist right now. Maybe next year ARM will have something ready)
Atom looks mostly dead however, with Intel beginning to advertize 16W Xeons. Can two 8W Centrino Atoms compete against a 16W Xeon? I doubt it. Intel only has the resources to choose one "hero" chipset in server-space, and it looks like it will be the Xeon again.
Right now, there are professional 1st-world developers developing apps for all those "incompatible CPUs". So it's possible and even cost effective under the right circumstances.
Yet someone is currently being paid to haul away this stuff as electronic recycling.
Snapdragon S4 completes the Linpack benchmark with 460 MegaFlops: http://www.androidauthority.com/snapdragon-s4-pro-vs-exynos-.... Snapdragon is rumored to use ~3 to 5 Watts of power.
Consider a Xeon E5-2690: which gets you 347,000 MegaFlops of performance in 135 Watts. http://www.intel.com/content/www/us/en/benchmarks/server/xeo...
Xeon just blows the ARM chip out of the water in Performance per Watt used. A modern day Xeon gives you somewhere on the order of 20x more performance per watt compared to a Snapdragon S4 Pro. Otherwise, all the electricity that it takes to run your supercomputer built out of crap chips will make your supercomputer more costly to run in the long term.
So my bet on memory is that Xeon wins that too.
On the multi-processor communications issue... I doubt that the Snapdragon would have PCIe connections... let alone a high-bandwidth low-latency connection like the Xeon's QPI.
There's a reason why supercomputers stick with Intel or AMD. They've got the processor power, the performance per watts, and I/O girth to communicate to other processors. A small embedded chip (like Snapdragon), just doesn't have the features to make a supercomputer out of.
Calxeda for example has spent a lot of time trying to solve this interconnect issue. Each Calxeda card is composed of 4 CPUs. Every CPU has 8 10Gb (aka: 1GB) ethernet links, 2 8xPCI connections and more.
Its a relatively young project, but they're moving towards the right direction. There are rumors that Facebook is deploying a Calxeda server.
I should have left that job sooner.
I'm able to fire off pretty intensive MapReduce jobs on an Amazon Elastic MapReduce cluster with many nodes for a fraction of the price mentioned in the post (less than $100).
While I can imagine I could repurpose all my MapReduce/Hadoop code to run on a single box - especially since Amazon does offer several high-memory instances today - I would be loathe to.
The MapReduce framework provides a really nice framework that lets me horizontally scale out compute, rather than vertically, and that is really handy at terabyte-data volumes (data warehousing and large-data analytics.)
Basically, if your computation would never exceed memory on a single machine, then it is more processor efficient to code a more simple multi processing method and run on a single box than to code a map reduce on a cluster.
But what if you're not sure of the input data size? Processors are cheap. Engineers are expensive. Code the thing once for map reduce and you don't have to worry about making the transition later.
at least two analytics production clusters (at Microsoft and Yahoo) have median job input sizes under 14 GB, and 90% of jobs on a Facebook cluster have input sizes under 100 GB
The other point the authors make is that DRAM costs are following Moore's law and terabyte-data workloads should "soon" be cost-feasible on single servers with DRAM.
Machines have 8-64 cores these days -- you don't want to write multi-threaded code every time you want to do an analysis. So you can write MapReduce, but use a multicore framework instead of a cluster framework (hadoop).
The unfortunate thing is that there is no popular open source implementation of a multicore mapreduce, so people use Hadoop, which is wasteful on small data sets, as mentioned.
But the great part about it is that you will use the same application code for both. I fully expect in 5 years or so that people will be running multicore mapreduce jobs on 100 or 1000 core boxes.
Not popular, I'd agree, but I have had a lot of success with one-off Akka projects. My mappers and reducers are usually under 10 lines of Scala (more if I'm stuck writing Java, obviously).
edit: if i reduce the style to a single font, wf_segoe-ui_light, then it still appears that way. so i guess it's an issue with the font, not the browser.
http://imgur.com/a/HNH5H From top to bottom: Chrome 27/beta, Firefox 16, IE 10 on Windows 7 x64.
If these message-passing graph algorithms are representative of a "bad fit" to the parallel map-reduce model, I'd say a 12% penalty is not a bad price to pay at all in return for all the benefits of the parallel cluster in other cases.
http://wavii.files.wordpress.com/2011/12/hadoop_too_big.jpg?...
edit: "pass on" (thanks, marshray)