Samsung announces 'Shinebolt' HBM3E memory
anandtech.com
anandtech.com
So, having more channels is always nice, but this "more cowbells" approach doesn't bring in the performance you expect in some cases.
it is perhaps underappreciated that CPU cores have gotten a lot more "thready" even in the absence of SMT, since each core has hundreds of instructions in flight. you know, kind of like how GPU cores (CUs, SMs) do...
I have a similar simulation code, which can saturate the cores pretty efficiently (It was running with >1.7M integrations/sec/core and scaling linearly on last decade's hardware).
But when you access to square submatrices inside a 3000x3000 matrix, and CPU prefetcher brings the whole n rows, and you discard most of the rows to get another part of the matrix, that bandwidth is wasted.
Instead, you can rearrange your matrices to make the prefetcher happy, and do not throw what prefetcher brings, because you make it bring which is going to be used next.
When I was doing last analysis runs, it was running with ~20% cache trash meaning I'm wasting 20% of the data I'm pulling in regardless. This is a huge loss, considering I was saturating memory bandwidth before the cores.
If I change how I fill and lay my matrices out, I can align my data with how prefetcher assumes, and way less data will be trashed during the process.
I didn't do that because it needed a codebase-wide refactorization, and my code was already 30x faster than reference implementation with better accuracy, so we didn't pursue that avenue.
However, my knowledge is limited on sparse side. If you want, I might try to dig into it a bit.
But complex physics like chemical reactions, particle tracking, any user field functions, is probably what is tricky to implement on a GPU.
A recent Microsoft paper proposed that the "perfect" inference machine is a ton of large-ish, modestly clocked, SRAM heavy chips piped together on a motherboard.
There is no RAM. The model weights are split up between the chips, and operations are pipelined like they already are in llms that are split into "layers". And there is no expensive interposer, or crazy interconnect like Cerebras. The chips are just wired together on a cheap motherboard, as the bandwidth between layers is relatively modest.
...It seems like this would work for training too. But what do I know? If it doesn't, the next "ideal" architecture is probably Cerebras's wafer-scale machines. And I am thinking they make use of TSMC's TSV tech (like AMD's X3D chips) to cram more SRAM onto a single wafer.
Theres also some fuss asserting that doing a bunch of matrix multiplication is an inefficient way to emulate physical neural networks, so maybe future architectures will be radically different?
Look over the course of history: there's often a roughly-fixed ratio between cpu performance & RAM bandwidth (+ size), and cpu/memory <-> graphics interconnect speed. Any of those factors get too far out of wack, and you've got a machine with poor bang/$.
GPU integration is old now. Ok I've read there's a mismatch between silicon processes used for logic (like cpu/gpu etc) and memory? But hey, this is 2023. Chiplets? Stacked dies? Or just do the best with available tech? It's not like there are no ICs that integrate RAM & compute.
With the right architecture, the "memory wall" could be (mostly) a thing of the past. With enough cores & on-chip RAM bandwidth, maybe specialized gpu processing wouldn't even be needed anymore? For example:
https://en.wikipedia.org/wiki/Massively_parallel_processor_a...
doing just that (LUT at the row buffer) https://youtu.be/9t1FJQ6nNw4;
doing parallel sense in NAND for computing AND between rows (which they call pages, but in DRAM they're called rows. It's the same concept, though.) https://youtu.be/QWL2Rw_VhHg
"Demystifying the Characteristics of High Bandwidth Memory for Real-Time Systems" https://upcommons.upc.edu/bitstream/handle/2117/362010/HBM.p...
> For this CAS latency, HBM (Observation 6) offers a significant advantage over regular DRAMs. This is because HBM has tCCD = 1 or 2 compared to tCCD >= 4 for DRAM. Accordingly, HBM can reduce the CAS latency component to at least half of its DRAM value.
Getting over 128gb of usable memory changes use cases for me, like running nodes on networks I wouldn’t have bothered with
Though apparently Apple applied for a patent for a hybrid DDR/HBM system, so maybe in M3/4/5? https://patents.google.com/patent/US10573368B2/