Raspberry Pi 2 Cluster Assembly Tutorial
pocketcluster.wordpress.com
pocketcluster.wordpress.com
There are two different forms of parallel processing, one for multi-core on a single node (i.e. shared RAM/peripherals), and one for multi-node (i.e. networked servers), where it can be highly impractical to share data due to the interconnections (usually Ethernet/Infiniband).
In most modern clusters, you have to use both approaches. RPi 2 is nice since you have multiple cores, and the cluster shown here allows you to play with the other side of the equation, multi-node parallelism.
EDIT: I looked into the parallella a bit. Still can't tell how well it represents multi-node parallelism, but it claims to have "a fast on chip network within a distributed shared memory architecture". So theoretically, this can simulate the interconnection of a large scale cluster, but I have no real knowledge of this board.
Beyond that a computer running Xen and a bunch of VM's is far more convenient, useful, reliable and probably more powerful.
Towards the other end of the scale, for example, we have a rack full of 1RU machines we use to run FEA and other purposes.
If it's for educational purposes, you don't need any hardware at all, as there's no aspect of it that can't be fairly easily simulated in software.
It's a toy.
This is certainly a neat project (and perhaps provides a nice talking point), but is it more than that?
The 4-core ARMv7 chip in the Pi 2 is way faster than the chip in previous Pis, but it's not a speed demon by any means. And the network I/O (and all I/O) being tied down to a slowish USB 2.0 bus means you're limited to ~90 Mbps through LAN, or ~200 Mbps if you use a USB 3.0 network adapter.
Even trying a fast SSD plugged into one of the USB ports (rather than a comparatively slow 'Extreme Pro' microSD card), you won't get a huge performance boost because of limited I/O bandwidth.
See many, many more benchmarks on this wiki: https://github.com/geerlingguy/raspberry-pi-dramble/wiki
P.S. I have a cluster of six Pi 2s running http://www.pidramble.com/, and I'm actually considering moving more of my hobby/very-low-traffic sites over to it, since it's a sunk cost of $400 whereas many of these small sites are currently on $5 or $10/month servers.
See: http://www.midwesternmac.com/blogs/jeff-geerling/getting-gig...
this number seems too low, hmm ah, you mean the 'funny' 1.7GHz i7
>or ~200 Mbps if you use a USB 3.0 network adapter
how is usb 3.0 going to help usb 2.0 pee? I think you meant GigE usb adapter, but looking at your tests even that merely got close to wired 100mbit speed :(
>sunk cost of $400
coincidentally thats the cost of a used i7 laptop :) or a purpose build home server
All in all such cluster is a cool hobby/learning project, but I wouldnt run anything serious on it.
For practical compute workloads, it's not very interesting. A Cortex-A7 core @ 900MHz has a peak of 1.8 Gflops (single-precision; 450 Mflops for double). Assuming that you could actually achieve peak for the cluster as a whole (you can't for non-trivial problems because of communication overhead), that would give you 43 Gflops in single precision and 10.8 Gflops in double.
A single Haswell core clocked at a very conservative 2 GHz has a peak of 64 Gflops in single-precision and 32 Gflops in double, without communication overhead because it's a single core.
More importantly, the Cortex-A7 is a 32-bit machine, so working on big problems is painful (though it does support LPAE). There's also much better availability of highly-tuned compute libraries for x86 and GPUs. You're going to be writing your own compute kernels to get near peak on the cluster of Cortex-A7.
Having to port stuff to a different architecture (or even having to rebuild tools/packages you'd ordinarily have on x64 but not on ARM) is a sobering experience.
Anyway, the Oracle JDK (1.8) runs speedily enough on these, and you can run ElasticSearch or Hazelcast just fine -- and focus on optimising things.
Like the quote goes, "if you want to write fast software, do it on a slow computer"
https://learnaddict.com/2015/08/03/raspberry-pi-stack-a-plat...
Code, tutorial, parts list, etc. here: https://github.com/geerlingguy/raspberry-pi-dramble
http://www.datastax.com/dev/blog/32-node-raspberry-pi-cassan...
A better way is to develop a multi-continent cluster, which will more adequately stress things like race conditions and jitter. But that's some $60/mo to have enough hosted VM nodes worldwide, and RPis win on cheapness pretty quick there.