The Raspberry Pi as a Poor Man's Transputer
jacquesmattheij.com
jacquesmattheij.com
The INMOS T800-20 was 20mhz and did 10 MIPS. However, it had four links, each doing 20Mb/s, so it can communicate at about the same speed it process data.
By comparison, a RaspberryPi is ~700Mhz and 847MIPs, but proportionally IO is extremely starved. It barely is able to keep up with 100Mb ethernet which gives only slightly more bandwidth than the the Transputer had 30 years ago.
There are numerous cool things about the RaspberryPi, like the graphics, the community support for cool software, the HD camera option (and the company making an HDMI input adapter board to connect to the camera interface on the RPi). However, it strikes me as a terrible choice for clustering and I feel sad whenever I see someone talking about doing that.
A much better choice would be an ARM chip with built in gigabit ethernet. One possibility would be to find PogoPlugs. They aren't just a bare board, but they can often be found for under $20 with hardware gigabit ethernet, an 800mhz or 1.2ghz processor, and 128-512megs of RAM. Another option that is a little more expensive would be to use Biostar J1800NH with 2 gig RAM sticks, pico PSUs, booting from cheap flash drives, could probably be put together for $105 per node.
"I know my hobby X is not good enough, but if I got enough of them, it might be way better than things actually designed to do X!"
When it was first proposed to me about a year or two ago, to massively connect RPis, I quickly did the math and showed a simple top-of-the-line graphics card ( I did the numbers for a Radeon 7970 ) outperformed RPis at a factor of something like 10 to 1 in price/GFlops. That's without even considering interconnections.
I know Jacques Mattheij is some kind of folkhero here on HN, so I'll tread lightly, but I find this weak as it's years overdue and comes complete with a non-committal call to someone else doing the action- not a value in this community. He implies his time is worth more than other people's:
"Doing this will be [...] well worth it." "If anybody builds this let me know"
[1] http://www.amazon.com/Introduction-Parallel-Computing-Analys...
Edit: But basically your comment about interconnecting RPis with ethernet is completely true. It's a huge waste of time.
I've been doing quite a bit of work using GPU's for problems that they are well suited to and the performance is extremely impressive. But as soon as you need more decoupling for whatever reason between the individual instruction streams they are not a very effective solution.
The communications overhead made by the GP is a good point and I'll definitely look into that, the xmos stuff linked elsewhere in this thread is very interesting as well.
Whether or not I'm a folk hero is up for debate, but feel free to speak as though you would to anybody else, I have 0 privileges here and would find it strange if people were to act differently because of my visibility. When I'm wrong I'm wrong.
And as to your last point: I simply don't have the skills to route the board required, that's why this is a blog post without schematics or board layouts. That's not because 'my time is worth more than other people's', see my other projects as proof to the contrary.
As for this being years overdue, the RasPi compute module has been out for only a few weeks.
I'll have a really good look at the datasheet of the compute module to figure out what its exact limitations are, thanks for the pointers to alternatives.
[1] http://en.wikipedia.org/wiki/Atari_Transputer_Workstation
I thought it was cool, but the Tramiels had other ideas.
The folks that created the Transputer have created a similar embedded processor.
Shortly thereafter (possibly due to some other medical devices project overrunning badly), the company went under, and one of the other interns and I bought the contents of the building at auction in April/May.
We were fortunate enough to know that "Box 7 : Electronics" happened to contain a bunch of T800s. So we spent the next summer cold-calling Perihelion's customer list ('Box 27 : Papers') and giving people some great deals. Those were the days...
Communication over GPIO is a cool idea but ultimately pointless.
The question then is how to distribute the code across the links to the other nodes, the transputer runtime would do that for you back in the day.
Likely there will be a ton of hacking involved before you get to the stage where a transputer dev-kit was out-of-the-box 3 decades ago but this feels like it is within the realm of the possible.
I have no illusions that there would be a ready-to-run package available right now that would give you all that, clearly some assembly is required.
I have no illusions that there would be a ready-to-run package available right now that would give you all that, clearly some assembly is required."
You're in luck. All the hard work in putting together a transputer devkit has already been done. XMOS dev kit (e.g. https://www.xmos.com/products/xkits/starterkit ) + XC programming language (e.g http://www.iraqigeek.com/2014/01/02/an-introduction-to-xmos-... ). Adding more cores seems to be plug and play.
XMOS sell ARM (Cortex M3) + XMOS hybrids, if you wanted to ease yourself into it. One of the later Amiga (post Commodore) computers has an XMOS in it too (though I haven't followed Amiga news for a couple of years now, so don't know if it's still available).
I do know there's a decent level of interest in using Erlang on the Parallella: http://forums.parallella.org/viewtopic.php?f=14&t=69
I thought it was really cool that you could take a Transputer chip, add 5V and Gnd and a couple discretes and have a working computer. Add wires to other Transputers, maybe some DRAM, and you had a nice computing surface.
The instruction set was pretty neat; some very odd but interesting ideas. I don't know how you would get superscaler-class performance out of a stack machine, though.
I wish they'd done the same package-level tricks (+5V and a capacitor and go) to an ARM; that would have been lots more compelling. Occam was just too damned strange to catch on, and the C compiler I had for my T400 was one of the worst that I ever used.
[0] http://www.greenarraychips.com/
[1] http://www.greenarraychips.com/home/documents/budget.html
What emerges is a sort of calculus for describing parallel systems which takes the amount of compute per second a node can do, the amount of data that node can "consider" during its computation, and the time it takes for all of the nodes to "react." In this case a node considers a data structure when it evaluates its state (say the head node of a list, or the median element of an array) and is constrained from taking action by a threshold of assurance that the data it is considering it accurate. The reaction time is the time between making a change to the data and the time at which all elements in the cluster considering that data consider it true.
I used to describe that to people like those big mechanical train signs at stations with all the letters on a flip wheel. People look at the sign and consider it, it is stable and 'true', then a train leaves the station and all of the letters start flipping as the sign changes to show the new truth about what is happening. People with no knowledge are stuck waiting for the sign to settle down before they can go to their track, people with knowledge can be heading for their track but if the track used on their arriving train is changed they will be put in motion again. The length of time it takes to change the sign is equivalent to the time it takes for a cluster to react. And and train station cannot usefully serve trains faster than the passengers can figure out which train to be on, nor can a cluster usefully process structured data faster than the truth of the relationships in that data can be ascertained.
It all collapses down to Amdahl's law of course but along the way can help you figure out where the inefficiencies are going to crop up in the system.
Other than for fun and learning - what are the advantages of a fabric vs. something like a GPU system? From my experience with micros it seems there's so much power lost in the peripherals that it's inefficient to use anything with a small number of cores for massively parallel systems.
Also, there is no isolation between the nodes, it's all one memory space (that's what triggered this to begin with), computing fabrics have security implications because the nodes are isolated from each other.
But it strikes me that the Pi compute module has way more power than the iPSC did, and consumes far less electricity. Switched gigabit ethernet would likely suffice -- you wouldn't need something like InfiniBand. Not sure if you'd want an out-of-band management network for something like this, but it'd be nice to be able to monitor/restart it if the network stack gets wedged.
You'd want some sort of backplane to host the SODIMMs and route data lines. Packaging density would mean convection cooling probably wouldn't cut it, so a fan would be needed. How many could you fit into a 1U rack, hmmm?
> How many could you fit into a 1U rack, hmmm?
Lots :)
How fast you reckon a GPIO based bus would be? Especially if the rpis are required to do something more than just shuffle data around.
I was wondering this too, but I never took the time to find out what this guys numbers could translate to in terms of actual throughput.
> There was a thread a while ago about a little board that you could connect 4 ways to its neighbours (for the life of me I can’t find it…),
That might be: https://www.indiegogo.com/projects/pshdl-board
So while 5 off-the-shelf x86 computers on gbe network do not sound as cool as 30 RPi modules with custom backplane, they should perform significantly better while costing about the same.
Processing power available to the masses is intimately connected to mobile phones and tabs. RasPi's are only one small step removed from a girl or a guy wielding a soldering iron.