This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.
Those two combined already result in a huge reduction in bytes per mm2, so with the same wafer processing capacity you're producing far less byte of memory. Add to that a complicated chain of HBM-specific packaging steps, and you're now also losing a decent bunch of perfectly-fine dies because rather than putting it into DDR you tried making a HBM sandwich and screwed up.
Even if the memory cells are the same and have an absolutely identical yield, HBM will always end up having a significantly lower output. That's just the cost of stacking, but some people are willing to pay the per-gigabyte price penalty in return for the higher bandwidth.
Soldering the DRAM onto a PCB is such a reliable process that there is almost zero risk of defects and even if a defect occurs the damage is limited. If the DRAM is soldered onto a DIMM the risk of a defect on the non memory hardware is non-existent. If the DRAM is soldered straight onto an SBC or GPU, then the DRAM can be removed to save the precious SoC or GPU chips.
Meanwhile HBM is the ultimate nightmare scenario. You stack up to 16 DRAM wafers on top of each other. One defect and the whole stack is worthless and that was actually the easy part.
In stage two things get even worse. You now have your accelerator chip and you must place the HBM on that chip. E.g. Blackwell GB300 has eight HBM stacks and the accelerator chip has a bigger area than the HBM. You must get the packaging right eight times in a row or you have wasted not only the DRAM silicon, but also the accelerator silicon because HBM cannot be removed and defects are permanent.
The issue here isn't just the yield of the HBM (which is obviously lower if you have taller stacks) but rather the yield of the combined HBM-based product, which is why doesn't make sense to say it needs more area but it is completely correct to say that HBM leads to more silicon being consumed. Hence it doesn't make sense to talk about yield of the HBM itself, because it is always part of an integrated product.
I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?
These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That's what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.
[edit] The packaging needed to support HBM is also much more expensive too. If demand for current HBM applications tanks and the manufacturing lines need filled, then maybe. Currently, the price point would make it very very niche.
HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.
> These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5.
The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.
So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!
These are not normal requirements, other than for a GPU.
Consider the 'MMA N matrices' primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there's nothing preventing it from being 1024 or larger if we have more memory bandwidth.
All this with a single instruction.
We'll have to figure out how to read and right to the surface of a black hole to get speeds that high.
If you have infinite memory bandwidth you just move the bottleneck back to compute so both have to grow simultaneously in lockstep.
What you should have said is that CPUs have so much compute headroom for matrix vector multiplication that simply adding more memory bandwidth would make them faster so every improvement in memory bandwidth is welcome.
The "move the bottleneck back to compute" bit is changing rapidly though. The first time a major hardware company ships a PIM chip, you can push for a few orders of magnitude more data through without being compute bound.
But in reality we also already have unified memory architecture systems, integrated graphics etc.
And memory is already expensive. It's downright hard to even get it though - you frequently would prefer not what's cheapest, but whatever is in largest scale production.
The only way I see HBM becoming competitive for consumer CPUs is if they solve the yield issues. Or if AI crashes so hard that people are putting those GPUs on fire sale and salvaging mass quantities of HBM off of them.
I’m assuming that it’s mounted in the same unit as the GPU, not in a handy-dandy socket.
It could be cracked open and extracted no doubt but if it’s glued in that’s going to be nigh on impossible to extract.
I had been pinning my hopes on HBM coming onto the second hand markets after there three years or so of use, perhaps I’m wrong.
Or 256-512 bits on medium to high end consumer CPUs if you're apple.
At least DDR6 is probably widening things 50%.
The advantage in going wide is transferring cache lines rapidly, not the CPU bus interface.
LPDDR6 PIM would primarily help the low end and mid range accelerator market. E.g. embedded models running on SBCs can be up to 1 GiB in size with acceptable performance, small models at 8 GiB become really easy to run at reasonable speeds on a smartphone and PIM enabled laptops or mini PCs make it possible to run 32 GiB models locally without compromise.
Of course this also assumes that the associated accelerators (NPUs) will catch up too, but the general point is that you won't need a 5090 or a 4090 anymore.
Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.
Running 17 chrome tabs doesn't benefit at all from that HBM and all the additional hardware+software complexities that come with it. You want a specialized coprocessor to handle specialized workloads. The GPU exists separately from the CPU for a reason.
You make HBM instead of DDR and the number of bits you're making goes down.
It's really that simple.
People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.
No, Mac laptops use LPDDR, currently LPDDR5X.
They have managed to pull this sort of thing off many many times. https://en.wikipedia.org/wiki/Reality_distortion_field
If I get 10% more performance for 50% more cost it really depends on one's needs, for example.
This distinction doesn't change what the performance numbers look like today, but it does inform what changes would be necessary for those numbers to look different tomorrow. E.g. Apple Silicon isn't fundamentally orders of magnitude more efficient than x86, they just used smaller features. Newer Intel and AMD chips made on equivalent processes _also_ get similar efficiency gains.
And honestly, they have historically had different markets.
When the design is for only one customer, you don't need to generalize things, and those things you generalize to give different customers different options has costs.
AMD will soon be a larger customer for TSMC than Apple (NVIDIA is already there) so Apple's pre-booking new processes is likely to be gone in the near future.
It's the address space that's unified, not always the physical hardware.
The data movement (when needed) is handled transparently in the background by page faults and other tricks.
Keep in mind that non-Pro threadripper is still only 256 bits wide and Pro is 512. And the memory is 30% slower than with an M5. So an M5 Ultra has 3x the memory bandwidth of the best threadripper.
This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.
So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.
The reality distortion is that people seem to believe it's HBM, or somehow it gives you extraordinary amounts of vram. Neither are really true.
DDR is optimized for latency and stability at the cost of bandwidth whilst GDDR is optimized for bandwidth at the cost of latency and stability. GDDR is pushed so hard these days that a small percentage of errors is expected and corrected because this is still faster than running it slower but more accurate.
GDDR7 often has 10-20x the total bandwidth but 3x the latency of DDR5. Graphical workloads want as much bandwidth as possible but care relatively little for latency. Conversely, applications love low latency but don't really see any performance benefit from higher bandwidth.
So basicallyt you have workloads that are diametrically opposed and running unified memory forces you to compromise.