I wonder if that university still uses their PS3 based supercomputer cluster
I wonder if that university still uses their PS3 based supercomputer cluster
The analogy I gave to my friends at the time was working in a restaurant kitchen with a tiny stovetop and 2-dozen microwave ovens; does some things really fast, but only if you can cut them up into small pieces that are microwave-friendly.
The hardware design was based on 1) This is currently the only way to hit that clock rate at this cost 2) Hyperthreading effectively halve all latencies 3) It's a fixed platform, the compiler should know exactly what to do. 3 was pretty laughable. It is only technically possible if you have huge linear code blocks without branching or dynamic addressing.
I would remind junior programmers that the PS2 ran at 300Mhz and had a 50 cycle memory latency and the PS3 ran at 3Ghz but had a 500 cycle latency. So, if you are missing cache, your PS3 game runs like it's on a PS2.
On the other hand, a lot of people overreacted to the manual DMA situation of SPU programming. DMAs from main mem into SPU mem had a latency of.... 500 cycles! Once people put 2 and 2 together, SPU programming became less scary. Still a pain in the ass. And, a lot of work to reach peak performance. But, more approachable for sub-optimal tasks.
Also, IIRC, shift by variable amount was microcoded and would take a cycle for each bit distance shifted.
Wait, it was possible to ship a CPU in 2006 without a barrel shifter? ARM1 was from 1985!
If it was something silly like "only constant bit shifts use the barrel shifter", I'm surprised that compilers didn't compile variable shifts as a jump table to a bunch of constant shift instructions... :)
I was absolutely floored too.
And it made very little sense to me either. PowerPC has probably my favorite main ISA bit manipulation instructions out there: rwlimi
https://www.ibm.com/docs/en/aix/7.2?topic=is-rlwimi-rlimi-ro...
Any logical shift, rotate, bit field extract (by constant) and more all in one single cycle instruction that's been included since the earliest POWER days. They had a barrel rotater in the core for that instruction.
The only thing that makes any sense to me is that somehow it would have been too expensive to rig up another register file read port to that sh input, so they just pump it as many times needed with sh fixed to 1. They seemed to be on some gate count crusade that might have payed off if they were able to clock it faster at the end of the day. It took the industry a bit to figure out that ubiquitous 10Ghz chips weren't going to happen, and the hardest lessons would have been right in that design cycle. : \
> If it was something silly like "only constant bit shifts use the barrel shifter", I'm surprised that compilers didn't compile variable shifts as a jump table to a bunch of constant shift instructions... :)
Variable shift isn't the most common op in the world, so as far as I know it was just listed as something to avoid if you're writing tight loops.
I think the main issue with the Cell design was that it was too "middle road" and wasn't specialised enough in either direction.
Eh, only if you count every SIMD lane as a separate "core" like GPU manufacturer marketing does. More realistically, you should count what NVIDIA calls SMs, where the numbers are more comparable (GeForce RTX 3080 has 80, for example).
PS4 - 18 GCN CUs, each has four 16-ways SIMDs for 72 SIMDs but each is 4 times wider so PS4 has the same number of ways as in 288 4-ways SPUs.
So there is not much difference imho and the GP is correct.
What do you mean? You can hide latency on SPU by double-buffering the DMA but in a shader there is no infrastructure at all and no way to hide unlike SPU, you just block until the memory fetch completes before you need the data.
> they are outright programmer-hostile
Depends on the programmer I guess, I enjoyed programming SPUs, don't know personally anybody who had complaints. Only read about the "insanely hard to program PS3" on the internet and wonder "who are those people?". It's especially funny because the RSX was a pitiful piece of crap with crappy tooling [+] from NVidia yet nobody complaining about SPUs mentions that.
[+] Not an exaggeration. For example, the Cg compiler would produce different code if you +0/*1 random scalars in your shader and not necessarily slower code too! So one of the release steps was bruteforcing this to shave off few clocks from the shaders.
Almost certainly not. A typical lifetime of a supercomputer is around 5 years, give or take. After that the electricity they consume makes it not worth continuing to run them vs. buying a new one.
See e.g. the Cell-based Roadrunner, in use 2008-2013: https://en.wikipedia.org/wiki/Roadrunner_(supercomputer)
ORNL's Titan lasted about 7 years: 2012 through 2019. Its predecessor, Jaguar, was 2005 to 2012. Also 7 years.
(At a previous job, we had a cluster that was about 15 years old. Of course, it had been expanded and upgraded over the years, so I'm not sure anything was left of the original. Maybe some racks and power cables.. :) )
[1] https://www.newsweek.com/here-comes-playstation-2-156589
I've heard on the grapevine that the PS3's OtherOS facility was internally thought of as another go at the same idea. "Look, judge, it's a general purpose computer for reals this time. Your own universities are using it in super computing clusters, without ever launching a game".
In my lab I found a couple of PS3s lying around several years ago, that hadn't been used in quite a while. (One of them may or may not have been adopted for less scientific purposes …)
>The Condor Cluster project began four years ago, when PlayStation consoles cost about $400 each. At the same time, comparable technology would have cost about $10,000 per unit. Overall, the PS3s for the supercomputer's core cost about $2 million. According to AFRL Director of High Power Computing Mark Barnell, that cost is about 5-10% of the cost of an equivalent system built with off-the-shelf computer parts.
>Another advantage of the PS3-based supercomputer is its energy efficiency: it consumes just 10% of the power of comparable supercomputers.
I wonder how significant the cost and energy savings were by the time the project was finished, and how long the cluster was actually used.
[0]https://phys.org/news/2010-12-air-playstation-3s-supercomput...
Actually, I just did a quick Google on this and it was "Sony" themselves that appeared to mention this! [1]
[1] https://www.tomshardware.com/news/cell-broadband-engine-ps3-...
An NVIDIA A100 GPU does about 20 teraflops. So you only need 25 of those chips to match the theoretical rating, and they have many other advantages like much higher memory per core, etc.
Also the PS3 apparently drew up to 200w, so a cluster that size would have drawn 352 kW.
This is way out of my field so I don't know the whole implications, but my understanding is Nvidia cards cam only reach these speeds at the loss of precision or full functionality, so it's an apples to oranges comparison versus non-nvidia chips.
NVIDIA had not even released the API for writing vectorised C code yet.
Hence all those 2007 supercomputer stories. They actually had a genuine use, because the Cell was fully implemented into the Linux kernel by IBM.
These guys did incredibly well for their time, and they were entirely right about vectorised code.
They were just superseded by the longer term trend of tying together multiple silicon dies and ASICs for HPC.