Not exactly weird, but on the PowerPC side "eieio", is my most favourite name for an instruction on any architecture - without fail, every time I come across it, I start humming Old McDonald!
Not exactly weird, but on the PowerPC side "eieio", is my most favourite name for an instruction on any architecture - without fail, every time I come across it, I start humming Old McDonald!
Sounds a lot like the PowerPC's rlwinm instruction.
The PowerPC 600 series, part 5: Rotates and shifts
https://devblogs.microsoft.com/oldnewthing/20180810-00/?p=99...
“?.” The experienced user will know what is wrong.
Macro: int EGREGIOUS “You really blew it this time.” You did what?
Macro: int EIEIO “Computer bought the farm.” Go home and have a glass of warm, dairy-fresh milk.
Macro: int EGRATUITOUS “Gratuitous error.” This error code has no purpose.> EGREGIOUS
and
> EGRATITUOUS
and... lots of stuff ^^
But the SPU ISA docs are available in many places, just googled and found this mirror: http://www.ece.stonybrook.edu/~midor/ESE545/SPU_ISA_v12.pdf - the various shifts start on page 117, but they're simpler than I remember - shift left in various forms, rotate (right) and mask (i.e. shift right) and rotate (right).
But, even looking at the instructions, it's not immediately obvious that every instruction is operating on a 128-bit vector - e.g. shl operates on 4 32-bit fields at once,and shlh operates on 8 16-bit fields at once, and so e.g. a single instruction can even shift by different amounts. Almost all instructions operate this way, only a few only use the first field in a vector, e.g. conditional branches (brz, brnz, brhz, brhnz). The other rotate instructions are interesting - you can do bitwise rotate across the entire 128-bit vector, rotate by bytes (very useful for memory accesses) and the wonder instruction shufb (shuffle bytes, see page 116) which allows each byte of the result to be pulled from a specific byte in of 2 vectors, or a constant (0,0x80,0xFF) which allows you do do all manner of things for example byte order swapping or different values together.
So, you have this really powerful instruction set that's vectorised on almost every instruction, and then you have a massive 128 registers per core coupled with 32KB local cache memory which is single cycle access. It's an absolute joy to program for on anything which will fit into the 32KB local cache (which includes your program). The only downside is that anything that doesn't fit into 32KB requires you to start a DMA request to copy a chunk to or from main memory, and wait for it to complete and/or continue doing other work in the background. With suitable workloads operating on 10KB or less, many people used this in a true double-buffered fashion, so the next set of data was transferring while you were processing the previous set, but this made lots of tasks quite complicated and many developers found it too hard to make good use of the SPU, and so it was often sadly under-utilised. One interesting effect though, is that DMA'ing your buffers took a similar time to a cache miss on the PowerPC cores, so actually even a naive approach that always triggered a small DMA from main memory in parallel with other operations often wouldn't see much latency and might still even be quicker than running the same task on the PowerPC cores where there was less opportunity for doing other work during a cache miss.
Another interesting feature of this CPU is that you can do a pre-emptive branch prediction so that the instruction cache is already full when a branch is takenm meaning that a branch can execute in a single cycle. Coupled with indirect branch (i.e. using a register value), you can give a CPU a branch hint in advance of the end of a loop as to whether it'll terminate or not and also avoid stalls. In some cases this could outperform the x86 branch prediction that was common at the time, which was just maintaining a cache of the last target of recent branches.
Overall, SPU programming was tough to master, but once you got something working the performance was phenomenal. I think that's why I liked it so much - because the sense of accomplishment was so tangible once you'd battled through it and emerged the other side with some amazing code. But nowadays, the power it provided, that was unheard of at the time (7 cores at 3GHz), isn't anything special over a mid-range i9.
SPU local store is 256kb, not 32kb. It's not single cycle access either, it's multiple cycle latency (6, if I recall correctly); it's still wicked fast though, it's about the same as L1 cache access on the PPU.
You're right about the latency too, but because it was a pipelined architecture all instructions had some latency, so while 6 sounds high, it wasn't really significant. Now I'm thinking about it (memories of re-arranging instructions manually, before I shifted from assembler directly to using GCC and intrinsics and letting the compiler worry about the interleaving), I also realise I'd forgotten the odd/even cycle split where you would pair instructions of different types (roughly ALU and non-ALU) together so they could execute concurrently.