What’s Motivating the Matrix Engine Movement in HPC?
nextplatform.com
nextplatform.com
On the one hand, I know this is a sentiment that I’ve seen echoed amongst some of my colleagues. Simulation codes just don’t FIT, sometimes, into the neat little box that these mature hardware coprocessors try to put them in. It’s easy enough for ad revenue peeps to adjust their data and models to a new architecture, since the only will they’re obeying is their own and the computational reality is more-or-less whatever they want it to be. With simulations a lot of these assumptions go out the door, because now you can’t just disobey fundamental laws when convenient. If you need to go through a calculation on a single core because you have to evaluate every state sequentially, sometimes that’s all that can be done.
On the other hand, this article seems to be subtly beating the hardware co-design drum, which in my opinion isn’t always a good idea either. One argument could be made that limiting ourselves to one computational approach actually encourages creativity because it encourages some bright young researcher to come up with a new way of looking at things to make the software better fit the hardware. An argument could also be made that co-design is sometimes a bad idea because it might result in things being shoehorned in when they have no place. I’m certainly guilty of experimenting with FPGA implementations which, at the end of the day took so long to compile they would never be useful.
As in all things, I think in this case both “sides” have a little something to offer.
To be fair, NVIDIA is also guilty of this. Their marketing is highlighting their tensor core units, claiming order-of-magnitude speedups on deep learning workloads, when in practice, workloads that involve convolutions are actually memory bandwidth bound and will not see a significant benefit.
HMC was interesting: the RAM operations could complete out-of-order (if one section of RAM was refreshing, maybe it'd be faster to out-of-order the responses). Too bad it never became popular. HBM / HBM2 still needs to account for refreshes, and still takes place in-order.
The Phi with 16 GB of memory onboard would probably make a very interesting developer workstation. Nothing will make code become more parallel faster than making developers plow their fields with chickens instead of oxen. ;-)
(forgive-me Seymour)
Threadripper only has ~100GB/s, best case, across its 4x memory channels. (25.6GB/s at 3200 MT/s). Upgrade to Threadripper Pro or EPYC (8x memory channels) and you're still only ~200 GB/s to main memory bandwidth.
--------
I'm sure Threadripper has more compute (and cache) than Xeon Phi, but a lot of that is 22nm process vs 7nm process.
https://versus.com/en/amd-ryzen-threadripper-3990x-vs-intel-...
https://insidehpc.com/2016/01/mcdram/
This one with MCDRAM (aka: HMC RAM from Micron) was well known to hit 400GBps bandwidth.
It does seem like this model is 14nm, as you've said. So I stand corrected on the process node for Xeon Phi: 14nm is more realistic representation, at least for the HMC model I was thinking about.
[0] https://www.anandtech.com/show/16195/a-broadwell-retrospecti...
Depthwise convolutions are different though - they usually are bandwidth bound.
POWER9 made a big splash with Summit, and POWER9 is no bandwidth slouch either. POWER9 also had Talos II so that "normies" could actually get a machine to play with under $10k.
Raptor Computing (makers of Talos) haven't said anything about POWER10 support yet. So its hard to get excited about a product that seems impossible to use.
IBM needs to understand that if there's no easy onboarding to a platform, it becomes a legacy one, as new things are built elsewhere.
You're absolutely right. 10K is a huge amount of money to spend on a computer, even a workstation. Most startups can't afford that, so they will never run POWER hardware, and the startups of today are the big corporations of tomorrow. POWER will likely disappear just like "minicomputers" did.
This makes me think of Apple making their ecosystem increasingly more locked in and restrictive. The result is that Apple devices become increasingly less fun for developers to play with. The end result is that more developers will migrate towards Linux. These people will then be less likely to develop software for Apple.
They'll stay with the IBM i, but there isn't much AIX can do that Linux on Xeon can't do as well as. AIX is not like their mainframe business (and even there, they have surrendered to Linux).
> These people will then be less likely to develop software for Apple.
They still have a pretty good UX for developers and entry-level machines. I'll probably replace my aging Mac Mini with an ARM-based one next year or so. What will really upset me is if MacPorts is no longer supported on ARM. I hope Apple is sponsoring someone ensuring key parts of the 3rd party ecosystem are there or Macs will become second-class developer machines.
NVidia NVLink is 300GBps GPU-to-GPU communications. In contrast, AMD EPYC is 50GBps CCX-to-CCX communications. Once again, there's no comparison with bandwidth, the GPUs simply win.
GPU-to-GPU is wider than CPU-to-CPU. VRAM-to-GPU is wider than RAM-to-CPU.
-------
If you have a problem that cares more about bandwidth than latency, then GPUs are simply the best architecture. Period. The pure compute performance is paired with absurdly huge bandwidth numbers.
Most consumer problems are latency-bound. HPC however, is specifically programmed to be bandwidth-bound instead. It takes effort and optimization to get there however.
Is that even true? I'd argue a lot of what happens on consumer devices can (or could) be done in parallel. There's some need for a fast response, but things like, responding to a button click and updating some internal state, take a trivial amount of computation to implement. Updating your UI in response to that button click, that's again something that can be parallelized.
blah = blah->next
A simple linked-list traversal: 50-nanoseconds on a typical consumer CPU. This could represent a DOM-tree traversal. Or it could be a Javascript object being allocated. Or it could be the text getting added to the end of your rope-data structure. Or it could be a binary-search (finding a new address to check next). Or a malloc() (which traverses a linked-list to find memory). This dereference happens very often.50-nanoseconds to hit main-RAM, 10-nanoseconds if L3 cache, 4-nanoseconds if in L2 cache, 1-nanosecond if in L1 cache.
This sort of operation is very fast on CPUs, especially if you're cached. Heck, your CPU won't even wait on the result: its going to branch-predict the answer and start executing the next bit of code before it even knows its correct.
----------
That same memory-dereference is 300-nanoseconds from GPU VRAM, and maybe 100-nanoseconds if you're in GPU L1 cache.
If you got a tree, its probably going to be far faster to traverse on a CPU. Where the GPU wins is Raytracing: you can have millions of individual tasks traversing the same tree, in parallel, at larger overall bandwidths than a CPU ever could.
IE: GPUs can perform 1000 binary searches faster than a CPU can perform 1000 binary searches.
But how many applications need 1000 binary searches? Usually, you need 1 binary search, and once you get the result, you need to do something else. (And that result can be branch predicted for even further accelerations to speed).
> GPUs can perform 1000 binary searches faster than a CPU can perform 1000 binary searches.
With several vector lanes and ILP on a 64+ core node, the 1000 parallel searches on the CPU may well beat the GPU. It'd have to be tested in any case.
You're right, 1000 is a small enough number where the CPU might still win. If its a large-enough data-structure to hit VRAM (ie: a binary search across a 8GB vector), that's what'd be best for the GPU.
Raytracing is millions rays per scene. 4k is 8,294,400 rays at 1-sample per pixel, and each one of those rays has to traverse a BVH tree (similar to a binary tree). Each of those rays MIGHT set off a 2nd or 3rd ray, depending on what its bouncing against. The big BVH tree that organizes all of the video game triangles / objects is shared for all rays.
That's the "1000s of binary tree searches in parallel" I was trying to reference. But I guess its closer to 10s of millions.
A64FX is HBM2 at equivalent bandwidth to GPUs (with lower power). CCX is much finer granularity than an entire GPU, so not a direct comparison. L3 bandwidth on EPYC is multi-TB/s.
Fat GPU nodes can readily overload network interfaces so if bisection bandwidth is your concern, CPU nodes are good.
> HPC however, is specifically programmed to be bandwidth-bound instead.
This is wishful thinking. Lots of applications used to justify the US exascale program (and others) are latency-bound. Climate, weather, unsteady CFD, and much of mesoscale materials science and molecular dynamics are run at their latency limit in most scientific studies (one-off scaling studies notwithstanding). There's an unfortunate disconnect between what scientific computing actually needs versus what funders and the media portray.
But only when using SVE512 SIMD-units, which are grossly similar to GPU SIMD units. At a minimum, SIMD reigns supreme. Even Intel only gets its max bandwidth when using AVX512 units.
Once you start rewriting your inner loops to run with SIMD, its not too difficult to start thinking about a dedicated SIMD-accelerator, or GPU, to do the job.
> L3 bandwidth on EPYC is multi-TB/s.
32-bytes / infinity fabric cycle. 16-CCX per EPYC chip. 3GHz == 1.5 TBps. I dunno about "multi-TB/s", but its over 1 TBps... sure.
But only if used in parallel. EPYC only has 32MBs of L3 per CCX. To achieve the full bandwidth, you need to split the problem into each CCX (which asserts a MESI-like "exclusive" lock when writing to a L3 location: preventing other L3 caches from reading-or-writing there).
Even then, EPYC's L3 cache is small compared to GPU-VRAM. Radeon VII 1TBps VRAM applies at full speed with atomics / synchronizations (in fact, the atomics / synchronization to L2 cache. I just don't have L2 numbers for that GPU...). I would expect Radeon VII's L2 cache to have more bandwidth than EPYC's L3 cache (Indeed: the Radeon VII L2 cache is in front of a 1TBps HBM2 cluster).
If we traverse up the GPU cache structure, you get 10TBps+ on __shared__ or LDS RAM, which is used as atomic-synchronization points or thread-barriers within a workgroup (a batch of up to 1024 cudaThreads).
As such: synchronization between threads (within a large workgroup), or even across the device, is reasonably efficient. The downside of this comparison is that 1024 cudaThreads only have access to 64kB of __shared__ RAM. So it isn't really comparable from a size perspective.
------
Your "TBps" estimate on EPYC's L3 cache however, is misleading. Because you spend significant amounts of MESI messages passing cache lines back and forth between CCX to get there. Compared to a unified L2 on the GPU (or unified VRAM), its just not really comparable.
EPYC L3 is somewhere between GPU L2 and GPU __shared__, in terms of memory hierarchy and complexity of use.
Regarding L3 bandwidth on EPYC, I have PDE solvers that exceed V100 global memory bandwidth for problem sizes that fit in EPYC L3 cache. I use NPS4 and eschew sharing data structures between CCX. This is the sweet spot for a sizable fraction of apps. The V100 is better value if you have more patience (and are thus able to batch or run larger problem sizes per device).
Hmm, I can certainly agree to that. There's some tricks to try to minimize that issue, but they're complicated to use and really screw up the overall architecture.
Batching up a larger batch of things to process can help for that. But GPUs have much smaller caches in general and are clearly designed for running out of VRAM. (With caches mainly for atomic / synchronization here and there). A100 only has 40MBs of L2 cache (last level), but the extra threads really eat up that space faster than you expect.
Fugaku has two really important features: Stacks of HBM on die, and the Scalable Vector Extensions. HBM is used by gpus as well. This gives A64FX comparable raw bandwidth even if the programming model is different. The SVE design is more general purpose than other short vector extensions. It abstracts over lane width / number of lanes at the binary level, and has some bypassing features similar to the old Cray machines that again have the effect of delivering massive bandwidth.
The result of this is that A64FX is competitive with GPUs on most HPC workloads, but with a single chip, single programming architecture, and lower overall power consumption. So I'd interpret the author's comments in that context. It's not so much about adding SIMD as a bad idea, but the specifics of how that needs to be implemented, and whether a dedicated coprocessor with a wholly different programming model makes sense vs well designed vector extensions backed by a high bandwidth memory subsystem.