Telum II at Hot Chips 2024: Mainframe with a Unique Caching Strategy
chipsandcheese.com
chipsandcheese.com
This is an amazing 23-minute video by the Microsoft programmer who developed the Windows NT Task Manager, among other things. He visits IBM and talks to engineers about the Telum chip architecture (Hot Chips 2023), used in the z16 mainframe. Special attention is paid to the cache.
In an episode on hard drives, he talked about how drivers for hard drives still report a constant number of sectors per track, so they must have a physical layout that matches that. Hard drive manufacturers are open about the actual layout of their drives and that they virtualize the hard drive for the OS so that it behaves well.
A few Microsoft engineers also dispute a lot of the facts of his stories about the development of the start menu.
Caveat emptor.
afaict (and I've worked with mainframes for a couple of years) this is spot on. poor signal/noise ratio but the facts are right.
[1] https://chipsandcheese.com/2024/09/05/an-interview-with-susa...
If someone is really into high performance, it's ideal to never have to wait for DRAM, either with predictive fetches or explicit cache warming. For that, the more cache you have, the better.
https://en.wikipedia.org/wiki/Transaction_Processing_Facilit...
I believe Unisys still makes x86-based mainframes running MCP.
My personal take:
The typical x86[1] is a sports car. Gets going fast, reaches most destinations fast, not great for driving for several hours, and not great at moving lots of cargo.
A mainframe is a freight train. Somewhat slow to get going, but can haul large amounts of cargo without breaks for a long time.
Mainframes weren't built for an interactive, highly variable, query-response workload; they were built for the classic overnight/monthly batch job that streams through a large amount of data.
[1]: It's not about the CPU, it's about the architecture around it, like this article talks about cache, expanded to I/O etc concerns.
This is an impressively fast design, but you'll get much more x86 for the same money -> x86 wins when you can scale out.
You might be confusing it with the AS/400 CISC ISA, which exists as an emulation layer on top of POWER, since all IBMi machines are almost identical to their POWER counterparts.
The AS/400 / 'i' are descendants of the System/38 and implement a "Technology Independent Machine Interface". Applications target this high-level interface, rather than the underlying hardware. Before first run (or when they're installed?) applications get compiled from abstract Machine Interface code to native code.
The software side seems to be more a tale of dichotomy. The MVS lineage is technically impressive but undoubtedly bizarre and old feeling. The TPF lineage seems like eventually somewhere the cloud movement will dip for certain cases so it is ahead of time. Linux is neither stale nor avant-garde, I guess that is their strategy to remain "contemporary". VM was always the most delightful one but internally forever the odd one out.
If mainframes were not competitive at that, they would have ceased to exist a long time ago.
Yes there's a lot of cache. But rather than try to have a bunch of cores reading each cache (sharing 96MB L3 for AMD's consumer cores), now there's a lot of separate 36MB L2 caches.
(And yes, then again, some fancy protocols to create a virtual L3 cache from these L2 caches. But less cache heirarchy & more like networking. It still seems beautifully simpler in many ways to me!)
From https://chipsandcheese.com/2023/03/12/a-peek-at-sapphire-rap... “ the chip appears to be set up to expose all four chiplets as a monolithic entity, with a single large L3 instance. Interconnect optimization gets harder when you have to connect more nodes, and SPR is a showcase of this. Intel’s mesh has to connect 56 cores with 56 L3 slices.”
There are probably design reuse and RAS considerations that make it not currently worthwhile to i.e. have a distinct physical design for SAP or whatever cores.
https://pages.cs.wisc.edu/~remzi/Classes/838/Fall2001/Papers...
The innovation here seems to be adaptive sizing so if by whatever algorithm/metric a remote core is idle, it can volunteer cache to L4.
Presumably the interconnect is much richer than contemporary processors in typical IBM fashion and they can do all the control at a very low level (hw state machinesµcoding) so it is fast and transparent. It will be interesting to hear how it works in practice and if POWER12 gets a similar feature since it shares a lot of R&D.
Have a look at the section "Cache setup" at https://chipsandcheese.com/2024/08/14/amds-ryzen-9950x-zen-5... for some real-world latency values. Once we're talking about a 100+ MB working set (i.e. DDR5 instead of cache), a top-of-the-line Ryzen 9950X has an access latency of about 100 ns. There is also some older data for a wider variety of CPUs at https://chipsandcheese.com/memory-latency-data/ - and there the older IBM z15 is in a class of its own.