Also, the latency with GDDR7 is pretty terrible. It uses PAM3 signaling with a cursed packet encoding scheme. At least they were nice enough to add in a static data scrambler this time around! The lack of RLL was kind of a pain in GDDR6.
Their secret is that the memory is manufactured within the chip package.
I remember Apple used to show slides depicting the M1 SoC as one unit containing a CPU, GPU, Neural Engine, cache, and DRAM all together. But slides shown at an Apple event definitely qualify for artistic license.
https://en.wikipedia.org/wiki/Apple_M1#/media/File:Mac_Mini_...
You could call that on-package, but it's not on-package in the same way that GPU HBM is, with that the main die and memory dies are packaged together on the same substrate. That's a much more difficult and expensive process, apparently packaging is the main bottleneck for H100/H200 production.
You can consider it like a small PCB that has the CPU die and the memory soldered very closely nearby (and like mentioned +memory channels)
That the memory is on a PCB close to the CPU/GPU certainly helps with signal integrity, but it is not by any means relevant here. The Apple platform has high memory bandwidth compared to x86 PCs because the CPU has a wide memory bus. You can get similar memory bandwidth out of high-end Epyc and Xeon CPUs which use standard DIMMs but with many more memory channels than a regular desktop computer.
[1]: https://www.nextplatform.com/2024/02/27/he-who-can-pay-top-d...