Not really? Apple is efficient because they ship moderately large GPUs manufactured on TSMC hardware. Their NPU hardware is more or less entirely ignored and their GPUs are using the same shader-based compute that Intel and AMD rely on. It's not efficient because Apple does anything different with their hardware like Nvidia does, it's efficient because they're simply using denser silicon than most opponents.
Apple does make efficient chips, but AI is so much of an afterthought that I wouldn't consider them any more efficient than Intel or AMD.
I wonder if a little cluster of Mac Minis is a good option for running concurrent LLM agents, or a single Mac Studio is still preferable?
On the higher end, building a machine with 6 to 8 24GB GPUs such as RTX 3090s would be comparable in cost (as well as available memory) to a high-end Mac Studio, and would be at least an order of magnitude faster at inference. Yes, it's going to use an order of magnitude more power as well, but what you probably should care about here is W/token which is in the same ballpark.
Apple silicon is a reasonable solution for inference only if you need the most amount of memory possible, you don't care about absolute performance, and you're unwilling to deal with a multi-GPU setup.
https://www.apple.com/mac-studio/specs/
Edit: since my reply you have edited your comment to mention the Studio, but the fact remains that the M2 Max has at least ~40% greater bandwidth than the number you quoted as an example.
One more edit: I'd also like to point out that memory bandwidth is important, but not sufficient for fast inference. My entire point here is that Apple silicon does have high memory bandwidth for sure, but for inference it's very much held back by the relative slowness of the GPU compared with dedicated nVidia/AMD cards.
It's definitely not what you'd want for your data center, but for home tinkering it has a very clear niche.
Is it? This is very subjective. The Mac Studio would not be "fast enough" for me on even a 70b model, not necessarily because its output is slow, but because the prompt evaluation speed is quite bad. See [0] for example numbers; on Llama 3 70B at Q4_K_M quantization, it takes an M2 Ultra with 192GB about 8.5 seconds just to evaluate a 1024-token prompt. A machine with 6 3090s (which would likely come in cheaper than the Mac Studio) is over 6 times faster at prompt parsing.
A 120b model is likely going to be something like 1.5-2x slower at prompt evaluation, rendering it pretty much unusable (again, for me).
[0] https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inferen...
The M4 Pro in the Mini has a bandwidth of 273 GB/s, which is probably less appealing. But I wonder how it'd compare cost-wise and performance-wise, with several Minis in a little cluster, each running a small LLM and exchanging messages. This could be interesting for a local agent architecture.
Apple's chips have the advantage of being able to be specced out with tons of RAM, but performance isn't going to be in the same ballpark of even fairly old Nvidia chips.
In any case how are you going to fit 50+GB in two (theoretically 24+24 GB) Nvidia cards without swapping to disk when the Mac in question has 64GB (also theoretically) available?
> In any case how are you going to fit 50+GB in two (theoretically 24+24 GB) Nvidia cards
No one was talking about only two cards.
That post read like a joke comparison and still does. Can you elaborate how it is relevant?
The parent of my initial comment in this thread said: "For inference, Apple chips are great due to a high memory bandwidth... It's a cost effective option if you need a lot of memory plus a high bandwidth."
My post was attempting to explain at a high level how 1) Apple SoCs do not really have high memory bandwidth compared to a cluster of GPUs, and 2) you can actually build that cluster of GPUs for the same cost or cheaper than a loaded Mac Studio, and it will drastically outperform the Mac.
If you want specifics on how to build such a GPU cluster, you can search for "ROMED8-2T 3090" for some examples.
I hope this helps.
https://machinelearning.apple.com/research/neural-engine-tra...
Until Apple can bang something out as good as an h100, it's no competition.
Cuda thrives bc of the hardware offering too.
Is the whole "unified" RAM a reason that the iMac and Mini are capped at 32G?
Unified memory is much more useful when you can get more bandwidth to it.
GPU cores are generally identical between the iGPUs and the discrete GPUs. Adding a PCIe bus (high latency and low bandwidth) and having a separate memory pool doesn't create new opportunities for optimization.
On the other hand having unified memory creates optimization opportunities, but even just making memcpy a noop can be useful as well.
The performance dependency on DGPUs doesn't come from the existance of a PCIe bus and partitioned memory, but from the fact that the software running on the DGPU is written for a system with high bandwidth memory like GDDR6X or HBM. It creates opportunities for optimization the same way as hardware properties and machine balances tend to, the software gets written, benchmarked and optimized against hardware with certain kinds of performance properties and constraints (like here compute/bandwidth balance and memory capacity, and whether CPU & GPU have shared memory).
I hadn’t used a PC in so long, I still thought that bios setting decided the division. TIL.
Lucky we have Asahi Lina to clarify the details.
e.g. My Ryzen iGPU reserves 2GB/32GB for itself (which Windows can't see) via BIOS and use 9 more as shared "unified" memory.
The base model is 256 gb. You can see it here:
Apples segmentation takes place in the nand firmware. The firmware contains the location in the storage configuration. And this may or may not be rewriteable. Iboffrc has done a video explaining how some of it works. It's all from reverse engineering though.
The top of the line is also not where Apple is gouging the worst. It's in the middle tiers that are actually relevant to many more people. Most don't have a need for 4+ TB main drives but 1-2 TB is a size that's pretty easy to justify for a lot of people and Apple's price is the only option for them and they're absolutely lining their pockets with cash at the expense of anyone not going for the bargain bin basic tier that can't hold 2 modern games.
It's worse than that -- 4TB gen 4 drives can be had for well under $300, sometimes $225-250, and that's for buying a drive outright, not "trading up" from a 256GB device. I think it'd be more accurate to say that you can get double the capacity for a _quarter_ of the price.
I also, as a side note, try to give the loosest most favorable (to my opposite) comparison because when I err on my side it becomes a "well actually" debate a lot of the time about how it's "not quite X times as many it's more like X-1 (so I'm not even going to touch that X-1 is still quite bad)" that is really tedious and annoying especially when the favorable version of the comparison is still quite bad for their point/side.
Except Apple wanted $3,000 for 7TB of SSD (considering the sticker price came with a baseline of 1TB).
I bought a 4xM.2 card and 4x2TB Samsung Pro SSDs, cost me $1,300, I got to keep the 1TB "system" SSD, and was faster, at 6.8GBps versus the system drive at 5.5.
Similar with memory. OWC literally sells the same memory as Apple (same manufacturer, same specifications. Apple also wanted $3,000 for 160GB of memory (going from 32 to 192). I paid $1,000.
Alternatively, you can get one of these[1] external Other World Computing NVME SSDs for USD1,190 right now. And then you can easily move all your files from your laptop to your desktop when you get home.
[1] https://eshop.macsales.com/item/OWC/US4EXP1MT08/ (15% off list price as of writing)
I'm considering getting one and a nice big monitor or TV. It needs to run x-plane 12 at decent speeds and maybe support a bit of light gaming. My macbook M1 pro is actually pretty decent for this but the screen is too small for me to easily read the instruments. I expect this will do better even in the base setup.
Otherwise my needs are pretty modest. I'd love to see steam add some emulation support for these things as I have some older games that I enjoy playing. I currently play those on a crappy old intel laptop running linux. I've also been eyeing a new AMD mini PC with the latest amd stuff (Beelink's SER9).
Seems pretty nice as well and seems like it is more performance for the money. Apple is doing its usual thing of charging you hundreds of euros for 50 euro upgrades. Get the base mac studio instead. It probably makes more sense if you are going down that path.
For example: https://www.amazon.com/dp/B08S47KBMC/
We really are living in the future if people are using these words in combination.
Though compared to this new mini a lot will feel clunky. Any HDD enclosure is certainly larger.
There is a reason for the popularity of those enclosure/hub combos that have the same footprint and color as the Mini.
I can't imagine anyone but Apple shareholders drooling at the taught of overpriced soldered memory would prefer a smaller Mac Mini case if ~0.5" more height would get you M.2 bays for storage.
[1] https://eshop.macsales.com/shop/external-drives/owc-ministac...
Upgrade your memory and connect it externally over USB-C. It works brilliantly
Fortunately I don't really see the point of using a mac mini, so this doesn't bother me too much, but... it's poor taste. You're holding it wrong was not cool the first time.
The issue is that Apple moved the storage controllers into their SOC. So they use raw nand chips, and you need to use ones that the SOC supports.
How do you love your internet speed compared to the internal stuff?
As a bonus, you can back up your computers and iDevices to the shared local storage instead of paying for (probably much slower to access) cloud storage.
Hell if I'm dumping cards from my camera or moving models across the network that's like 64-128GB max.
This basically proves that Apple shot themselves in the foot for AI on mobile by artificially restricting RAM for so long! Heck, even the Neural Engine has turned out to be basically useless despite all their grandstanding.
So alas, their prior greed has resulted in their most popular consumer iDevices being the least AI compatible devices in their lineup. They could’ve leapfrogged every other manufacturer with the largest AI compatible device userbase.
It's really crazy that some android phones at half the price give a better browsing experience than many iPhones, especially the non "Pro" ones.
Tbf it seems to be browsers that kill memory on any platform. The Web is now just mobile breakpoints that load needlessly high res images.
Of course, the browsers are problematic for memory but that's not new and hilariously the first iPhone was supposed to work only with web apps. It's not like if they couldn't spend a few dollars more on their expensive premium hardware to guarantee a good experience for all use cases.
This is the big problem with Apple today, the milking at every single step, the stupid pricing ladder and scrooge attitude is extremely distasteful and completely unreasonable for the price asked.
What they shot was us. My 14 Pro won’t do AI despite having a better NPU than an M1, all because Apple chose - intentionally - to ship it with too little RAM. They knew AI was coming and they did this anyway.
Although having played with it on my MBP it’s clear I’m not missing much. But still.
And their npus weren't added in anticipation of LLMs, imo. You give em too much credit.
It's Unified RAM. So that memory is also used for the GPU & Neural Cores (which is for Apple Intelligence).
This is actually why companies moved away from the unified memory arch decades ago.
It'll be interesting to see as AI continues to advance, if Apple is forced to depart from their unified memory architecture due to growing GPU memory needs.
I don't understand - wouldn't the OS be able to do a better job of dynamically allocating memory between say GPU and CPU in real time based on instantaneous need as opposed to the buyer doing it one time while purchasing their machine? Apparently not, but I'm not sure what I'm missing.
The usual reasoning that people give for it being bad is: you share memory bandwidth between CPU and GPU, and many things are starved for memory access.
Apple’s approach is to stack the memory dies on top of the processor dies and connect them with a stupid-wide bus so that everything has enough bandwidth.
Today, the industry is moving toward unified memory. This trend includes not only Apple but also Intel, AMD with their APUs, and Qualcomm. Pretty much everyone.
To me, the benefits are clear:
- Reduced copying of large amounts of data between memory pools.
- Improved memory usage.
- Generally lower power consumption.
And besides, what Apple is doing is placing the RAM really close to the SoC, I think they are on the same package even, that was not the case on the PC either AFAIK?