What AMD Learned from Its Big Chiplet Push
spectrum.ieee.org
spectrum.ieee.org
Software is hard to change. How ironic, and yet of course completely true.
Most software is sold with “works with existing hardware” as a selling point.
Most vendors don't have the luxury of being able tailor both to eachother.
We saw the result with Intel Itanium: great concept, absolute failure due to software incompatibility. On the other hand, the PS3's Cell architecture demonstrated that innovative platforms aren't impossible given the right incentives.
On the contrary, even if you better not access memory too randomly to get full performance with GPU compute, today you can often access gigabytes of memory randomly IF needed (PS3 Cell had 256kb.. but you were double or triple-buffering for send/receive so effectively it was less) and even if it's better avoided it gives a whole lot of freedom in practice. Doing RTX style raytracing would've been doomed on the Cell.
Which are the same kind of features driving mesh shaders, or more recently, DirectX work graphs.
It doesn't have to be RTX style raytracing.
I took RTX style raytracing as my own current demoscene stuff is doing raytracing with just compute shaders (not needing actual RTX support)
- A flawed theoretical basis (super-smart compilers and VLIW CPU architectures were the key to faster general-purpose computing performance)
- Many, many years of design/development delays before Intel actually shipped working hardware
- Competing CPU architectures (x86, RISC, etc.) increased their performance enormously during Itanium's delay years
- When you finally could get Itanium systems with decent performance...they were damned expensive, and their bang-for-the-buck performance was mediocre at best.
Apple did it three times with the Mac 68K -> PPC -> x86 -> ARM
The same was true for x86 - by the time Apple started selling x86 machines, the x86's were so much faster than even G5's it was hard to notice any slowdown.
And, while I assume there are uses in which some x86 MacPros still perform better than their ARM counterparts (1.5 TB of RAM must help), those are far and in-between.
There were situations:
> The 603 was intended to be used for portable Apple Macintosh computers but could not run 68K emulation software with performance Apple considered adequate, due to the smaller processor caches. As a result, Apple chose to only use the 603 in its low-cost desktop Performa line.[12][13] This caused the delay of the Apple PowerBook 5300 and PowerBook Duo 2300, as Apple chose to wait for a processor revision. Apple's use of the 603 in the Performa 5200 line led to the processor getting a poor reputation.
- from https://en.wikipedia.org/wiki/PowerPC_600#PowerPC_603
(The original PowerPC 601 had a 32KB unified cache, the problematic 603 had separate 8KB I- and D-caches, and the "problem-fix" 603e had separate 16KB I- and D-caches. My recollection is that the 68K emulation software was developed for and on the 601, and relied on some clever lookup and/or jump tables fitting into that CPU's large - for the time - L1 cache.)
It’s a shame I didn’t have A/UX on the IIci though (and don’t have the IIci anymore) but Unix didn’t look like it was the future at the time.
The 6100/60 (PPC 601-60Mhz) was definitely slower with emulated applications.
Besides that, the 68K emulator didn’t emulate the floating point unit in 68040s making it incompatible with higher end applications.
Only the high end 8100/80 was always faster than most 68K macs under emulation. But still high end 68040s were faster running 68K Macs.
But this was largely alleviated with the third party SpeedDoubler extension that was a much better emulator.
This isn’t even to mention the awful emulator performance of the 603 Macs.
But this is the only benchmark I could find between PPC vs x86 for emulation.
https://barefeats.com/quad06.html
On a semi related note: Apple still doesn’t have an answer on the high end when it comes to competing with the fastest x86 PCs + GPUs
I think they decided not to pursue that market segment. Their current high-end is quite a lot of computer (the lack of support for dGPUs on the MacPro is disappointing, and while 192GB ought to be enough for anyone, it could sport an external memory controller and a DDR5 bus for those crazy people who really want terabytes of RAM).
When you go beyond that you enter the midrange tower server market and those machines don’t come cheap either. The day after the Pro’s release I priced a similarly equipped Dell (same number of cores, same memory and storage) and it was the same price (a bit less), but it could be expanded much further.
It was also much hotter, louder, and uglier.
> Most vendors don't have the luxury of being able tailor both to eachother.
Microsoft has to care about niche vertical market apps that will never be updated.
But compared to the Xbox, PS3 "won" at the tail end of the generation, I would credit this to the supperior exclusive titles and the sales outside US. In Europe, Japan an South America, Playstation sells a lot more than Xbox. Where I live (Brazil) almost no one buys Xbox. As an anecdote, in 20 years, I met only one person who owned one (a X360) and several who had Playstations. And the X360 owner only had it because she worked at Microsoft and got it as a gift from the company.
I bought a slim as soon as they came out as a BluRay and media player. It was cheaper than about 95% of the other offerings at the time, had a way better UI and remote, and wasn't going to get locked out when content keys were rolled. It supported upnp media playback, so I could stick a mediatomb machine in a closet and use a nice remote and UI in the living room. A side benefit being that you can get dirt cheap laser modules and keep it going forever.
What open source software does is to turn hardware architecture into a commodity.
Itanium and Alpha, it seems, were a bit too early to this party.
I.e., to the extent they reduce the need for source code to cater to the underlying CPU, less effort is needed to port software to a new CPU.
I think GCC and especially LLVM are really reducing the barriers to bringing new CPUs to market.
Maybe for GPUs as well because of shader compilers?
And there are lots of temporary solutions in programming.
Once, we get a reasonable lean and modular software stack for most usage contexts, we can start to deal with new hardware models with their specific software optimizations.
A big mistake from AMD software teams is their worshipping of c++... they should have sticked to plain and simple C (not the gcc/iso extented one), actually I'll one step higher: plain and simple assembly for their hardware.
It also makes writing big old messes easier than C.
Writing simple, elegant and efficient code is hard. It takes "polish", which usually means not stopping the moment as you have "something that works", going back, refactoring, shaking out those TODO comments...questioning all the assumptions and abstractions you added earlier, etc etc.
It is less worse to write and try to keep sane plain and simple C, that to be dependent on a c++ compiler, not to mention that c++ code is no more immune from toxic code than C. Actually, this is the other way around: since the syntax is beyond-reasonable complex, c++ toxic code is way worse than c toxic code.
C is a lesser evil than c++.
OpenMP has all these process pinning environment variables that not enough people bother playing with. Or we could all start writing MPI programs. IMO is is helpful when the ecosystem forces you to explicitly think about communication.
Also the universe is constantly telling us that we can have as much bandwidth as we want, but latency is hard and coupled with distance in ways that it won’t let us violate, so we should play along and do more NUMA. Listen to the universe.
Ok the last paragraph just might be madness on my part.
As someone who regularly deals with NUMA in high speed networking, I wish the opposite.
It's shocking how much crossing a NUMA node just guts performance. 10Gbps becomes 5Gbps. 100Gbps becomes 20Gbps. Your application is too big for a single NUMA node? Sorry, your $80k investment only bought quarter of a fast computer, not a whole fast computer. The other three quarters are useless.
CPUs with wide fast interconnects between nodes and no locality to PCIe bus would make this situation much better. If AMD can do it then they'll walk all over Intel in some markets.
Modern mobile phones essentially have everything except memory and storage in a single chip, the Northbridge died a decade ago, and recent Zen chips can even function without a Southbridge.
Does this mean the processors have built-in USB and PCI controllers? Or is this more a case of "if you don't need many peripherals, you can do without a southbridge"?
I suspect the reason they can do without the southbridge is because many high-speed connections are PCIe these days (including storage and networking). But I don't see a good reason why Zen would have on-chip USB or RS232 controllers.
Socket AM5 has an integrated x16 PCIe bus, three x4 PCIe buses (one usually used for the chipset, one for NVMe, and one for USB4 ports - but other assignments are of course possible), a dedicated DisplayPort connection, a three-port hybrid USB3/DisplayPort interface, a dedicated USB3 connection, SMBus & I3C, and a HDA/SoundWire connection.
The AM5 Southbridge is essentially just a PCIe hub combined with some PCIe-to-USB and PCIe-to-SATA interfaces. If you don't need those, you can just leave it out! In practice you'll virtually always have it for the SATA ports and extra PCIe lanes, though.
Don't forget that USB is also a high-speed connection these days. 10Gbps is very common, and USB4 goes up to 40Gbps in Gen3 - with 80Gbps for Gen4 on the way. RS232 is indeed not included, but that's pretty much a dead port anyways: no reason to use it for onboard peripherals, and consumer motherboards rarely have a header for it these days.
AM5, I don't know.
i.e. CPUs need to be huge to achieve their performance targets, but wafer defects make for poor margins with chips that large, so they split up into chiplets to allow mixing and matching of good parts.
But SOCs are nowhere near that limit yet, so their tiny dies are growing to bring more more functionality on-die and reap the benefits of integration.
I wonder if that's it; simply keeping a PCIe link awake between various chips uses a bit of power to drive the trace capacitance, for instance.
Intel already launched a processor with 16GB on-package MCDRAM in 2016 (Knight's Landing Xeon Phi). You can even buy an Intel Xeon with 64GB HBM2 today. Nvidia likewise has been packaging HBM with their server GPUs.
Embedded DRAM (eDRAM) been used for long time in the mobile and console space. e.g. IBM's POWER7 (e.g. Nintendo Gamecube) and Intel Haswell products. However, using a logic process node to make DRAM cells is wasteful. Packaging technologies have advanced sufficiently that you now regularly see regular DRAM dies (LPDDR, HBM) being put on-package.
But all of that is packaging and manufacturing technologies. We're still taking to DRAM over a memory bus like we're still living in the '80s. The true innovation I'm looking out for is for a company to stick its neck out and use a different communication standard to talk with the DRAM modules. Something like the CLX.mem standard, which is used in the server space to talk to memory expansion modules.
The amount of people that need memory speeds and latencies beyond what SO-DIMM and CAMM can handle, but only need 16GB of RAM is absolutely tiny.
Is there a good overview of how much of a benefit the onchip ram is?
It’s all volatile storage with different uses.