Raspberry Pi 2 Cluster Case, Part 2
pocketcluster.wordpress.com
pocketcluster.wordpress.com
I really wish there were single board RPi / ODROID like devices available that sported support for a really performant and efficient interconnect fabric... like RapidI/O or something similar.
Not in the price range of RPi but much more powerful. You can also try the Beaglebone Black... Also, from the ODROID block diagram, the Ethernet PHY is connected directly to the CPU MAC, afaik the issue is the bad (cost?) choice of the Raspberry Pi...
I don't know of any actual switched-fabric mezzanine cards available for the Parallela. If you have a link I'd be very interested. A similar approach would be to exploit a mini PCI-Express connector like the one on the Jetson. The Jetson also has a daughterboard connector available, but my understanding is that it's proprietary and I don't think NVIDIA offers a switched-fabric card.
I own a couple of ODROID boards and while the network performance is better than what you'll see on the RPi, it's still nothing to write home about. And that's sorta my whole point... if you're working with large enough problem using these low power boards that you've got to offload compute work to multiple devices, the interconnect fabric has got to be much better than what is commonly available right now.
However, I'm fairly convinced that future systems are not going to be powerful monolithic systems but instead large collections of low-power cooperative compute units. It's not a huge stretch to claim that these exercises are worthwhile efforts to learn how to best exploit future hardware architectures on hardware that exists and is available for purchase today.
I certainly would be happier, assuming limitations like the interconnect fabric were well solved, if a common unit of compute quanta was more along the lines of a €50-99, 10-15 watt, board rather than €1-2K, 1-1.5k watt, server.
The best way to fix this is to put the actual OS on an external disk. You absolutely must bootstrap from an SD card, there's no way around that. But that can be entirely read-only and then you bounce to the external drive. Or maybe you could do something like a PXE boot instead - copy a system image down from a server and boot that. Personal opinion, this is the only way I would try using Pis on a long-term basis.
The Jetson TK-1 board from NVIDIA includes a real SATA connector. Because it's GPGPU-capable it's also vastly more powerful than even a Pi2. It also has USB 3.0 so running swap on a memory stick would be much more performant.
You can still write by just doing "mount -o remount, rw /" beforehand.
I'm still sort of aggrieved that Solid State via DDR got caught up in lawsuits and is still commanding exploitative pricing. SATA for Solid State isn't exactly optimal either, though obviously not nearly as terrible as USB - SATA - SSD. SSDs with PCIe interfaces are so common now, it's shame not to see wider support for it.
A little tidbit I saw in HP's "The Machine" promotional talks was the idea of "fabric attached memory". If ARM 64 & RapidI/O ever make it into widespread use, it would be great if there were various forms of memory with RapidI/O interfaces commonly available.
Also, don't forget that USB is just another serial protocol, if it wasn't USB it would just be some other serial protocol. It may not be the best option, but with stuff like the rasspi I imagine price is a bigger concern.
Edit: found them https://www.scaleway.com/
https://fsdata.se/server/raspberry-pi-colocation/
Edit: They did that, apparently the service is "on pause".
http://raspberrycolocation.com/
For 36 Euro's (39.54 USD) per year, they will host your RPi with a 100Mbit pipe and 500GB of traffic.
That being said, for light web hosting purposes, a cluster of six Pi 2s performs roughly 70% as well as a similar cluster of 6 VMs on Digital Ocean. See more on this project's wiki, in the performance/benchmarking section: https://github.com/geerlingguy/raspberry-pi-dramble
https://rwmj.wordpress.com/2014/04/28/caseless-virtualizatio...
I've got so many other projects on the back-burner I shouldn't even be reading about this sort of thing. But it does look fun!
More RAM is always nice, of course.
For learning distributed computing, I think has some advantages over just running all of the nodes on a single powerful computer. It means you can't hide from doing things scalably -- i.e. network costs between your processes are real; if your load isn't well-distributed across your nodes the OS scheduler can't save you; etc.