The flip side of this is if the GPU can access the main system memory then I could see this being useful for loading big models with much more efficient "offloading" of layers. Even though bandwidth between GPU->LPDDR5 is going to be slow, it's still faster than what traditional PCI-E would allow.
The caveat here is that I imagine these machines are $$$ and enterprise only. If something like this was brought to the consumer market though I think it would be very enticing.
(If anybody from AMD is reading this, I feel like an architecture like this would be awesome to have. I would love to run Llama 3.1 405b at home and today I see zero path towards doing that for any "reasonable" amount of money (<$10k?).)
Edit: It's at the bottom of the article. These are designed to be meshed together via NVLink into one big cluster.
Makes sense. I'm really curious how the system RAM would be used in LLM training scenarios, or if these boxes are going to be used for totally different tasks that I have little context into.
- Branch prediction and speculative execution - Out of order execution - Massive physical register files and register renaming - Cache predictors - and many more I'm sure.
Speculative execution is the big one for me, just because of the information leakage possible through it. It's there because you'd have to pause fetching new instructions until the result of a conditional branch is known, which has knock-on effects to instruction scheduling... But how big are these effects? Do some certain combinations of features supercharge or work against each other?
I'm sure there's people looking at such things inside Intel and AMD, but it doesn't seem like there's much out there for public consumption.
Look at how atrocious the CPUs were in the PS4/Xbone generation for an example of this.