Apple M5 could ditch unified memory architecture for split CPU and GPU designs
notebookcheck.net
notebookcheck.net
Unified memory is the only reason Macs are so coveted right now for local AI. A single 192 gb ram Mac costs less than the equivalent in standalone GPUs.
What are the good use cases for very large memory amounts?
DeepSeek v3 for instance has 671B params, but should have the memory bandwidth of a 37B dense model with a batch size of one.
You could also try Qwen 2.5 32b, which you should just work with ollama or LM Studio with no config changes.
I've got a 32gb M1 Max and a 24gb 4090, and I barely ever run models on my Mac, as the memory bandwidth and compute for prefill is much better on the 4090. But I'm essentially locked out of Llama 3 70B class models, which I only use via API.
[1] See: https://www.reddit.com/r/LocalLLaMA/comments/186phti/m1m2m3_...
I also have a home server with 2x3090 and 2xA4000 (80GB vRAM) - yes it's a lot faster, but it's a pain in the ass to build, it takes up a lot of space, uses 10x the power, and honestly - cost about the same as my MacBook Pro.
I hope Apple sticks with the architecture. Even if its not very practical, its great to have it as possible.
The software will still see a single memory pool
I know im disagreeing with the article
This does not follow. Intel is shipping unified memory processors with CPU cores and GPU cores on separate chiplets but still sharing the same memory controller (on a third chiplet, for Meteor Lake and Arrow Lake). AMD is about to launch Strix Halo, a high-end mobile processor that is rumored to consist of one or two CPU chiplets and an IO die with a big GPU and 256-bit memory controller.
> Twenty-four x86-architecture ‘Zen 4’ cores in three chiplets
> Six accelerated compute dies (XCDs) with 38 compute units (CUs), each with 32 KB of L1 cache, 4 MB L2 cache shared across CUs, and 256 MB AMD Infinity Cache™ shared between XCDs and CPUs
> 128 GB of HBM3 memory shared coherently between CPUs and GPUs with 5.3 TB/s on-package peak throughput
It's hard to rule out their ability to create silicon that is a step change.
I still want to try last year's Wizardry remake – which actually emulates the original Apple II code (or subsequent NES code), with a capability of displaying the original interface if desired.
The original XBox (2001) had 64MB. I think my PC from 1998 had that.
Edit: [2] The tweet doesn't even mention about UMA. The interpretation is entirely made up by Notebookcheck, I feel like I am reading WCCFtech again making stuff up.
I am just thinking if this allow Apple to do something crazy like 1024bit LPDDR5x or HBM3e memory solution.
[1] https://www.anandtech.com/show/21414/tsmcs-3d-stacked-soic-p...
[2] https://x.com/mingchikuo/status/1871185666362745227?ref_src=...
The actual rumour from Kuo is that they’d move to a chiplet style design where the CPU tile and GPU tile are independent. This is actually in the article as linked.
That does not however mean that unified memory would go away. It’s just a new packaging system.
Such a cool name! And it says just what it does.
https://spinoff.nasa.gov/node/8965
https://spinoff.nasa.gov/sites/default/files/thumbnail0000_2...
https://pubs.aip.org/asa/jasa/article/92/4_Supplement/2376/7...
Body Electric supported the Convolvotron for visually programming VR simulations with 3D sound:
https://news.ycombinator.com/item?id=24266722
Did you ever meet (or better yet get a tour of Ames from) the late Ron Reisman, and see the virtual reality, flight simulator, and air traffic control systems his research lab developed?
Vertical Motion Simulator:
https://www.youtube.com/watch?v=5-lHcv_olkE
Marvin Minsky flies a simulator and wears VR goggles:
I needed a username in the early 90s, I had just finished a paper where we microcoded a CM-2 to support high-throughput convolutions with spatially varying kernels for Hubble image correction (before the launched the eyeglasses mission), and I decided I could be the hero or anti-hero of convolution.
https://www.macintoshrepository.org/724-kpt-convolver-1-0
I can't find any demos of it on youtube, but it's the kind of obscure retro thing that LGR loves to review. He's really into the better known Kai Power Goo, which is a bit more accessible to kids than KPT Convolver:
As a couple of others have mentioned, smartphones/tablets/laptops seem to be the driving force in UMA's spread.
Is it really, though? It seems like almost every SoC small enough to be implemented as a single piece of monolithic silicon has gone the route of unified memory shared by the CPU and GPU.
NVIDIA's GH200 and GB200 are NUMA, but they put the CPU and GPU in separate packages and re-use the GPU silicon for GPU-only products. Among solutions that actually put the CPU and GPU chiplets in the same package, I think everyone has gone with a unified memory approach.
Much like dual socket servers, where each socket can address all memory, these new servers have two memory systems, one optimized for CPU and another optimized for GPU. Seems like a good idea to me, why serial/deserialize complex data structures between the CPU and GPU, which are then bulk transfered, and then checked for completion. With NUMA you can pass a pointer, caches help, everything is coherent, and it "just works". No more failures when you don't have enough memory for textures or a LLM, it would just gracefully page to the CPUs memory.
However the M4 Pro has 256 bits wide, M4 max 512 bits wide, and M2 Ultra has 1024 bits wide. GPU workloads are latency tolerant and embarrassingly parallel, don't see how allowing a CPU to make random accesses is going to hurt the GPU much.
The prices they charge just to go from 16GB to 32GB of RAM is outrageous ($400 for Macbook pro).
It should not cost that much! 2x Mac mini M4 16GB/256GB should not cost the same as 1x Mac mini M4 32GB/512GB!
Can someone help explain this in a way that isn’t just absolute price gauging of the higher end customer base? Are the components genuinely that much more expensive?
No, it's the same reason Nvidia has a vastly higher margin on datacenter cards:
It creates this weird dichotomy of having arguably the best value computer on the market in the base mac mini with 16gb of RAM and 256gb of storage and some of the absolutely worse value upgrades (like spending $400 on 16gb of RAM or $200 on 256gb of storage).
There's not much to explain here; they price gouge upgrades because they can. People that want/need MacOS for their work will pay for it, even if begrudgingly. I'm not necessarily happy about paying that much for these spec bumps but the benefits of using a Mac still outweigh the cons for me.
It's a pretty normal pricing strategy. It's more common than not. Most products or services you buy anywhere will be sold at higher margins for more premium offerings.
It might seem strange when compared to legacy PCs with socketed components, but this isn't that, nor are most products. Even among PCs this isn't strange anymore: go take a look at MS's pricing on their first-party PCs.
Calling this "price gouging" is not really the right use of the term -- usually it refers to price increases of basic necessities in emergency situations.
Either way, my point is that flat margin pricing is exceedingly rare. Everywhere from the grocery store to the car dealer is charging higher margins on more premium products.
Luxury cars have higher margins than economy cars. Organic milk has higher profit margins than regular milk. And Macs with 32GB of memory have higher profit margins than Macs with 16GB of RAM. The fact that the desktops PCs of our past priced RAM upgrades nearly at cost was an outlier; a courtesy, not anything normal.
https://en.wikipedia.org/wiki/Price_discrimination
It is basic microeconomics that a seller wants to be able to get as high of a price as buyers are willing to pay, but since different buyers have different abilities and willingnesses to pay, a seller can maximize their revenue by providing options at different price points.
Especially with societal wealth gaps, the people able and willing to pay higher prices are going to be able to pay higher price premiums, resulting in higher profit margins.
Price gouging, as a meaningful term, is restricted to:
https://en.wikipedia.org/wiki/Price_gouging
>Price gouging is a pejorative term used to refer to the practice of increasing the prices of goods, services, or commodities to a level much higher than is considered reasonable or fair by some. This commonly applies to price increases of basic necessities after natural disasters. Usually, this event occurs after a demand or supply shock.
Using the term "price gouging" anytime a potential buyer thinks a seller is asking for too much money renders it meaningless. I ask for as much money as the buyers for my labor will pay, as I assume the people selling to me do also.
It's just business, you try to earn as much as possible (and that could involve not maximizing in a specific transaction to incentivize repeat business in the future). But in no way is anyone under any duress when deciding to buy an Apple device, so if a buyer does not feel like being price gouged, they should buy something else.
Seems like they think Ultras aren't worth the investment, let alone building a true "unleashed" SIP.
Apple never says "hey what's the fastest and most powerful thing we can build for X price", they always box themselves in with space or energy constraints, so they never truly compete for the high end. The existing Mac Pro body was their chance to do that, and instead they put something designed for a smaller chassis in there.