Pico Cray – Small scale distributed computing
extremeelectronics.co.uk
extremeelectronics.co.uk
One way the RP2040 really punches above his price class is internal bandwidth. It has four AHB-Lite bus masters (two M0+ cores with one read-write port, and the DMA engine with a read and a write port) connected a full crossbar switch. There are six SRAM banks of which the four largest are commonly used word interleaved, but can if you want full control over memory timing and bandwidth you can use the uninterleaved alias instead. Most chips priced around a RP2040 (and the SPI flash to go with it) have only a single small SRAM bank. I've found the DMA engine surprisingly pleasant to work with (easy to understand, no fixed request routing to fight, but still very flexible which can save a lot of interrupts). Whoever designed the combination of the PIO channel I/O coprocessors and this DMA engine had a clever tinkers mind. I have to stop myself from wasting time looking for ways to save CPU cycles with trickery that only makes the code hardware to maintain.
One of the limiting factors is that RP2040 expects to boot from SPI Flash and there is no documented way for external SSRAM (the Synopsys DW IP that is used by RP2040 as SPI flash controller probably even can do that, but its documentation in RP2040 datasheet seems somewhat redacted).
(I would not be at all surprised if a single RPi were as powerful as twenty Cray-1s.)
Couldn't find any numbers for a Pi Pico.
0: https://web.eece.maine.edu/~vweaver/group/green_machines.htm... 1: https://en.wikipedia.org/wiki/Cray-1
As a rule of thumb, modern scalar pipelines can sustain one ALU op per cycle, and you see nearly all Linux-capable CPUs quoted in GHz. So we should expect gigaflops, minimum. And indeed [1] suggests that the Pi 4 is capable of 13.5GFLOP, so about 84 Cray 1's. (The further speedups come from the fact that ARM also has NEON vector instructions, and from multiple cores.)
The Pi Pico, on the other hand, does not have a floating point unit. So it emulates it in software (soft float). The C SDK docs [2] suggest 13.8kHz (!) operation for single-precision add. I'll be generous and suggest that the 2x cores could double this performance. So then, it'd achieve 0.0086% of the Cray-1's performance. Oof.
If you're willing to do integer arithmetic, things look much better for the Pico, of course - it runs at 125MHz and the above scalar rule of thumb applies.
[1]: https://web.eece.maine.edu/~vweaver/group/green_machines.htm...
[2]: https://datasheets.raspberrypi.com/pico/raspberry-pi-pico-c-...
Here is the homebrew cycle-accurate Cray-1, and it's a really impressive project:
You're correct this isn't really a Cray other than the vaguely rounded shape, but it is still adorable and looks super fun to program.
I've always wondered if a solution based of multi-master I2C support could work: all childs shout a unique ID and the parent answers with their assigned I2C address. Children that do not win the arbitration will detect collisions and stop because of arbitration errors. Unsure how well it could work with many children.
See https://www.i2c-bus.org/multimaster/ and https://www.i2c-bus.org/i2c-primer/clock-generation-stretchi... for introductions to I2C multimaster and arbitration. The spec is of course not open.
I'd be interested in discussing the topic more!
Always loved the CRAY aesthetic. Saw one in a museum a while back, couldn't get away from it.
In comparison to that model of Cray I'd be curious how the Raspberry's fair. I actually learned linux on a bunch of 486/586 boxen given to me by my neighbor, so powering a Pi2 from my laptop's USB port and getting a full linux distro really impressed me.