Microsoft is first to get HBM-juiced AMD CPUs
nextplatform.com
nextplatform.com
I run BareMetalSavings.com[0], a toy for ballpark-estimating bare-metal/cloud savings, and the things you can do with just a few servers today are pretty crazy.
What I wanted to get at is that the pure core count can be misleading if you care about power consumption. If you don't and just look at performance, the current CPU generations are monsters. But if you care about performance/Watt, the improvement isn't that large. The Zen1 CPU I was talking about had a TDP of 180 W. So you get 6x as many cores, but the power consumption increases by 2.7x.
Note that this doesn't intend to be used for accounting, but for estimating, and it's good at that. If anything, it's more favorable to the cloud (e.g, no egress costs).
If you're on the cloud right now and BMS shows you can save a lot of money, that's a good indicator to carefully research the subject.
With a quantization of it you can run larger contexts and go a bit faster. 1.4 tok/sec at 8b quant with offload to a 6GB laptop GPU.
Speculative decoding has been being added to lots of the runtimes recently and can give a 20-30% boost with a 1 billion weight model running the speculative token stream.
The Max variant is something they are using in their own datacenters. It would be possible that they would use an HBM solely for themselves, but it would be cheaper overall if they did the same thing for workstations.
Guess what the Ultra chips use? That’s right, a silicon interposer. :)
But the more I think about it, the more I bet they are creating a native Ultra chip that is not a combo of two Max chips.
I bet the Ultra will have the interconnect so you can put two together and get the often rumored Extreme chip.
They will have enough volume for their own datacenters, that the Mac Studio and Mac Pro will simply be consumer beneficiaries.
It makes more sense in this framing to put HBM on these chips. And no DDR5.
In this case, the M4 Max has neither HBM nor the interconnect. I’d love to see someone de-lid and get an X-ray die shot.
The M3 Max dropped the area for the interposer to connect two chips, and there was no resulting Ultra chip.
But the M1 Max and M2 Max both did.
I have yet to see an x-ray of the M4 Max to see if they have built in support for combining two, have used area for HBM or anything exotic, but they have done it before.
Could you recognize HBM support in an x-ray?
As for the Ultra, they used to have 2.5 TB/s of interprocessor bandwidth years ago based on M1, so I hope they would step that up a notch.
I don’t put much stock in the idea of the 4 or 8 way hydra. I think HBM would be more useful, but I’m just a rando on the interwebs.
I'd be really surprised if we see any consumer CPUs with HBM anytime soon. But it would be cool!
I’m not alone right? This article seems to be complete ai nonsense at various points confusing the gpu and cpu portions of the product and not at all giving clarity on which parts of the product have hbm memory.
By now I get that no one else cares and I should just stop coming here.
For anyone that read the article which product has hbm attached? The cpu or gpu? What is the name of this product?
There’s literally nothing specific in here and the article is rambling ai nonsense. The whole site is a machine gun of such articles.
I weep for the internet we had as children.
Be better.
[0] High Bandwidth Memory
The second is that you should never make assumptions about the audience of your writing, and their understanding of the topic; provide any and all information that might be pertinent for a non-subject-matter-specialist to understand, or at least find the information they need to understand.
They're almost certainly not on this forum, and they're not reading your post. So who is that quip directed at?
I don't know much about the site in the OP, but I work on the assumption that almost anyone could be reading comments on links to their site on this forum.
It's directed at them, you and even myself.
And it's directed at me? You don't know the first thing about me. PFO.
"Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting."
The point, in any case, is to avoid off-topic indignation about tangential things, even annoying ones.
I'd argue that I also provided value by solving the complaint I made by spelling out what it stood for, for those who might not know.