Playstation 3 Architecture
copetti.org
copetti.org
There was also that time we ran out of open file handles ahead of key demo at one of the tradeshows. I remember rewriting the whole audio streaming subsystem to piggyback on the seekfree texture loading system so that we could use the that single big file as a source for both streaming texture and audio data. Did that in 24 hours less than 2 days ahead of the show, fun times.
[Edit]
Another fun thing from that era was that almost no one other than the in-house Sony Studios used the SPUs well. X360 came out just ahead and had tooling that was miles ahead(PIX!) so everyone started from the X360. None of the workloads we're vectorized so it was absolutely brutal to get them to fit into the SPUs. For a long time a good chunk of that silicon was just idle. If you did start from the PS3 though because of the vectorized workloads you had insane cache coeheranch and usually ran better on the X360/PC as well.
[1] https://www.reddit.com/r/Games/comments/fnl0o1/why_playstati...
(ideal if you use accessibility tools, like text to speech, or want to read from an eBook or legacy browser)
If you spot a mistake, please log an issue on the repo (https://github.com/flipacholas/Architecture-of-consoles). Thanks!
I wonder if that university still uses their PS3 based supercomputer cluster
An NVIDIA A100 GPU does about 20 teraflops. So you only need 25 of those chips to match the theoretical rating, and they have many other advantages like much higher memory per core, etc.
Also the PS3 apparently drew up to 200w, so a cluster that size would have drawn 352 kW.
This is way out of my field so I don't know the whole implications, but my understanding is Nvidia cards cam only reach these speeds at the loss of precision or full functionality, so it's an apples to oranges comparison versus non-nvidia chips.
Almost certainly not. A typical lifetime of a supercomputer is around 5 years, give or take. After that the electricity they consume makes it not worth continuing to run them vs. buying a new one.
See e.g. the Cell-based Roadrunner, in use 2008-2013: https://en.wikipedia.org/wiki/Roadrunner_(supercomputer)
ORNL's Titan lasted about 7 years: 2012 through 2019. Its predecessor, Jaguar, was 2005 to 2012. Also 7 years.
(At a previous job, we had a cluster that was about 15 years old. Of course, it had been expanded and upgraded over the years, so I'm not sure anything was left of the original. Maybe some racks and power cables.. :) )
In my lab I found a couple of PS3s lying around several years ago, that hadn't been used in quite a while. (One of them may or may not have been adopted for less scientific purposes …)
The analogy I gave to my friends at the time was working in a restaurant kitchen with a tiny stovetop and 2-dozen microwave ovens; does some things really fast, but only if you can cut them up into small pieces that are microwave-friendly.
The hardware design was based on 1) This is currently the only way to hit that clock rate at this cost 2) Hyperthreading effectively halve all latencies 3) It's a fixed platform, the compiler should know exactly what to do. 3 was pretty laughable. It is only technically possible if you have huge linear code blocks without branching or dynamic addressing.
I would remind junior programmers that the PS2 ran at 300Mhz and had a 50 cycle memory latency and the PS3 ran at 3Ghz but had a 500 cycle latency. So, if you are missing cache, your PS3 game runs like it's on a PS2.
On the other hand, a lot of people overreacted to the manual DMA situation of SPU programming. DMAs from main mem into SPU mem had a latency of.... 500 cycles! Once people put 2 and 2 together, SPU programming became less scary. Still a pain in the ass. And, a lot of work to reach peak performance. But, more approachable for sub-optimal tasks.
Also, IIRC, shift by variable amount was microcoded and would take a cycle for each bit distance shifted.
Wait, it was possible to ship a CPU in 2006 without a barrel shifter? ARM1 was from 1985!
If it was something silly like "only constant bit shifts use the barrel shifter", I'm surprised that compilers didn't compile variable shifts as a jump table to a bunch of constant shift instructions... :)
I was absolutely floored too.
And it made very little sense to me either. PowerPC has probably my favorite main ISA bit manipulation instructions out there: rwlimi
https://www.ibm.com/docs/en/aix/7.2?topic=is-rlwimi-rlimi-ro...
Any logical shift, rotate, bit field extract (by constant) and more all in one single cycle instruction that's been included since the earliest POWER days. They had a barrel rotater in the core for that instruction.
The only thing that makes any sense to me is that somehow it would have been too expensive to rig up another register file read port to that sh input, so they just pump it as many times needed with sh fixed to 1. They seemed to be on some gate count crusade that might have payed off if they were able to clock it faster at the end of the day. It took the industry a bit to figure out that ubiquitous 10Ghz chips weren't going to happen, and the hardest lessons would have been right in that design cycle. : \
> If it was something silly like "only constant bit shifts use the barrel shifter", I'm surprised that compilers didn't compile variable shifts as a jump table to a bunch of constant shift instructions... :)
Variable shift isn't the most common op in the world, so as far as I know it was just listed as something to avoid if you're writing tight loops.
I think the main issue with the Cell design was that it was too "middle road" and wasn't specialised enough in either direction.
Eh, only if you count every SIMD lane as a separate "core" like GPU manufacturer marketing does. More realistically, you should count what NVIDIA calls SMs, where the numbers are more comparable (GeForce RTX 3080 has 80, for example).
PS4 - 18 GCN CUs, each has four 16-ways SIMDs for 72 SIMDs but each is 4 times wider so PS4 has the same number of ways as in 288 4-ways SPUs.
So there is not much difference imho and the GP is correct.
What do you mean? You can hide latency on SPU by double-buffering the DMA but in a shader there is no infrastructure at all and no way to hide unlike SPU, you just block until the memory fetch completes before you need the data.
> they are outright programmer-hostile
Depends on the programmer I guess, I enjoyed programming SPUs, don't know personally anybody who had complaints. Only read about the "insanely hard to program PS3" on the internet and wonder "who are those people?". It's especially funny because the RSX was a pitiful piece of crap with crappy tooling [+] from NVidia yet nobody complaining about SPUs mentions that.
[+] Not an exaggeration. For example, the Cg compiler would produce different code if you +0/*1 random scalars in your shader and not necessarily slower code too! So one of the release steps was bruteforcing this to shave off few clocks from the shaders.
[1] https://www.newsweek.com/here-comes-playstation-2-156589
I've heard on the grapevine that the PS3's OtherOS facility was internally thought of as another go at the same idea. "Look, judge, it's a general purpose computer for reals this time. Your own universities are using it in super computing clusters, without ever launching a game".
Actually, I just did a quick Google on this and it was "Sony" themselves that appeared to mention this! [1]
[1] https://www.tomshardware.com/news/cell-broadband-engine-ps3-...
>The Condor Cluster project began four years ago, when PlayStation consoles cost about $400 each. At the same time, comparable technology would have cost about $10,000 per unit. Overall, the PS3s for the supercomputer's core cost about $2 million. According to AFRL Director of High Power Computing Mark Barnell, that cost is about 5-10% of the cost of an equivalent system built with off-the-shelf computer parts.
>Another advantage of the PS3-based supercomputer is its energy efficiency: it consumes just 10% of the power of comparable supercomputers.
I wonder how significant the cost and energy savings were by the time the project was finished, and how long the cluster was actually used.
[0]https://phys.org/news/2010-12-air-playstation-3s-supercomput...
NVIDIA had not even released the API for writing vectorised C code yet.
Hence all those 2007 supercomputer stories. They actually had a genuine use, because the Cell was fully implemented into the Linux kernel by IBM.
These guys did incredibly well for their time, and they were entirely right about vectorised code.
They were just superseded by the longer term trend of tying together multiple silicon dies and ASICs for HPC.
> The accelerators included within PS3’s Cell are the Synergistic Processor Element (SPE). Cell includes eight of them, although one is physically disabled during manufacturing.
This was disabled in software (by syscon) early on in the boot process. Many people "unlocked" theirs without stability issues. My understanding was that it was a yield thing, and definitely not intended for general use ;)
> This makes you wonder if IBM/Sony/Toshiba hit a wall while trying to scale Cell further, so Sony had no option but to get help from a graphics company. Interviews from early 2nd party developers confirm this: https://www.ign.com/articles/2013/10/08/playstation-3-was-de...
That's also why the PS3 has two separate RAM, compared to the xbox 360's unified memory - Sony was trying to do that as well.
> HDMI connector
idk if anyone remembers this, but sony was talking about multi display gaming, such as having a status display. There was a prototype that had two hdmi ports and three ethernet - sony claimed it would also be a home server and router at that time. devkits did include two hdmi ports, and i suspect that the two screen claim was something that they made up almost on the spot, but who knows.
https://commons.wikimedia.org/wiki/File:PS3_e3_2005_prototyp...
Truncated URL. Was it copied from another comment?
Which is an excellent way to brand how much of a pain it was for Devs to wring performance out of it.
Some of the later PS3 games still look incredible. The biggest graphics limitation is resolution.
Then option (c) is only possible if you reduce the game to the lowest common denominator(with option option (a) becomes more appealing 2-3 years after the fact).
Nowadays the difference is less significantbetween PS4 and PS5 (or Xbox One and Xbox Series), but it's still a non-trivial amount to maintain; it would have been vastly more difficult with orders of magnitude in performance (between a PS2 and a PS3), or different programming paradigms (PS3 cell vs PS4 x86).
The PS1 was basically a repackaged project from their Nintendo partnership. They didn't really have time to develop proper development tools for it, and almost as a consequence of that, the development environment was very scrappy - it hooked into existing PCs and included a bunch of libraries that developers were somewhat familiar with using. As a result, the developer toolset was relatively easy to use. It allowed many developers to get started making 3D games and experiences very quickly. This caused developers to be swayed to the Playstation ecosystem early, and drove all 3D resources into Playstation development away from the Sega Saturn, which had the typically convoluted development environment from past generations (further exacerbated by Sega bolting together additional chips onto the Saturn to try to compete with the PS1).
PS2 came along, and Sony was already deviating from their easy to use console debut. Their "emotion engine" was notoriously hard to develop for, but Sega with the Dreamcast didn't have the legs to compete with Sony's momentum from the PS1, and quit the console business. Nintendo also screwed the pooch that generation with the Gamecube and Microsoft didn't seem to be a threat with the Xbox which was a major flop in Japan.
Along comes PS3, and Sony goes full speed ahead on their hubris - expensive console, impossible to develop for, an architecture somewhat reminiscent of Saturn's hodge podge of chips. At this point they completely lost their way from what made the PS1 set them up for multi-generations of success.
It's interesting to consider how the ecosystem matters when developing these hardware products, and how easy it seems for companies to lose sight of that.
This is what I remember from the Cell (both BE and the “serious” version IBM used in blade servers). It’s a fascinating machine, but being inconvenient to write software for is a key weakness.
Consoles are one of the last holdovers from the era where computers were designed with the hardware and software built and integrated from the bottom up into a cohesive package. Where they'll do a new one only every 7 years or so, and users just expect to have to buy new software for them every time. So it sounds like a space where they could get away with weird innovations and risks, but because of the need to keep cross-platform development feasible, it's not really an option.
I don't know. I see why it is the way it is, but I'd like to live in a world where consoles could get away with weirder things. Nintendo has kind of occupied that space with their touch screens and motion controls and whatnot, but the Switch is also more standard than ever before. Probably a net good for consumers though, with how much software is able to be ported to it.
The 3d was still being figured out in the 90s, which is what lead to so much diversity in product lineups. Not just for game consoles, but video cards, graphics API, etc. Being weird/unique was necessary because everyone was treading in uncharted territory.
Prior to the 3d era, game consoles had repurposed/customized off the shelf components. The 6502 and it's variants were used in all kinds of consoles and computers. And the 68000 was used in many, many more.
Gaming consoles have come full circle.
Same benefit it's always had: fun and innovation
[0]Yeah Apple says the GPU[1]is entirely their design at this point, but then quietly reached some kind of agreement with Imagination in the past year or so.
[1]Though considering how often GPU drivers need game specific updates and how fast GPUs change and improve, each GPU is almost like it’s own little bespoke platform.
The IBM PC was an exception and only because IBM couldn't prevent Compaq to carry on.
Ironically, the CEO of nVidia once stated that nVidia is a software company - the graphics cards are just dongles that monetize the software. This is what Sony missed.
The story I heard was that when sony finally went to buy a graphics card, they wouldn't pay for the software side of things. When nobody could get any performance out of it, they went back, cap in and and asked for the software to go with the hardware. Don't know if it's true, but I developed on PS3 and it feels true.
As I understand it, it was US developers and leaders like Mark Cerny that begged sony to stop fucking around and just make a console with a big cpu and a big gpu.
You can also see this in the early Cell Evaluation systems which at first had 6800s in SLI and switched to a 7800 GTX before the RSX was ready: http://www.edepot.com/playstation3.html#Early_PS3_Models
[0] https://en.wikipedia.org/wiki/Dennard_scaling#Breakdown_of_D...
They looked at the G80, but that was too early and risky to jump on at the time. So, they settled for the 7800.
It's a shame. If they had delayed the release and went with the G80, the PS3 would have crushed the 360 as far as graphics. Instead, a whole lot of Cell SPU time had to be dedicated to shoring up the PS3's GPU issues to bring it on par with 360 titles.
I want to say it's mentioned in The Race For A New Game Machine too, but I'm not 100% on that.