Llama.cpp can do 40 tok/s on M2 Max, 0% CPU usage, using all 38 GPU cores
twitter.com
twitter.com
Apple are particularly well placed to do this with the unified memory architecture of the M series processors.
Local LLMs are the future, not in the cloud, I really wouldn't be supprised if Apple drop something along these lines on Monday. I will be looking out for a "oh, just one more thing" after the headset demos.
I'm surprised they have been holding back on improving their GPUs if this is true. Was there any technical achievement in the recent years to enable this leap?
Honestly, Apple doesn't strike me as any better or worse off than their competitors. Maybe AMD, given how late to the party they are. Intel has AVX fallback acceleration through Haswell though, and libraries like OpenVINO offer extra acceleration paths for modern Intel CPUs. ARM has ARMnn for any multicore ARMv8 system, and of course Nvidia has CUDA support for multiple arches and OS. Projects like the ONNX Runtime are unifying all of these inferencing interfaces, giving developers multiple deployment options without resorting to walled garden implementations. WebGPU may obsolete them altogether.
Local LLMs are awesome, but Apple has some work to do if they want to lead the pack.
Over time, sure. Maybe even on the top-tier Macs as a demo. But I wouldn’t expect anything better than LLaMa.
I would, however, expect them to come to their senses regarding how inadequate Siri is.
I fully expect the Neural Engine inside Apple SoCs to take up 80% of the transistors in the future, up from ~10% today.
Citation needed. I don't think general public has any desire/care for that. Like they don't care about a locally run Siri.
On the other hand, Netflix did (does?) use AWS.
Consumers don't care where the LLMs are run, but if companies can off-board the computing costs onto the consumers, then they will save millions per month, but the other big limiter for local llms is storage and network costs.
> Local LLMs are the future
I think the issue with local-llms will be how can businesses talk to them? How do I use my local LLM with HN, Twitter, Shopify. Can my local LLM help me request a refund from a Shopify order or would Shopify have to fine-tune their own model? Then what? I download a 300GB file just so I can refund an order?
Clearly this isn’t an XOR proposition. Use web query to get relevant data (ie to generate dynamic prompt) then use local (ie private data) to calculate. Repeat ad nauseam.
Example: please get me the same vegetables I ordered last week from Whole Foods.
The incentive may actually be the other way around. If it has to be run in the cloud, it has to be a service. And selling a service is more attractive to companies (even Apple nowadays) than selling a piece of hardware.
They don't need to know how you solved that, but local LLMs are the way to solve it.
> Can my local LLM help me request a refund from a Shopify order or would Shopify have to fine-tune their own model? Then what? I download a 300GB file just so I can refund an order?
LLMs are like an implementation detail: Shopify wouldn't have a single "Shopify" model, but they might have models for specific things they need to do to generate value for themselves. Likewise you might have some local model that acts specifically like a personal assistant and is able to initiate a refund for you.
OpenAI's plugin spec already works for advertising capabilities to LLMs you don't own: https://server.shop.app/.well-known/ai-plugin.json
There's lots of use cases where you know the answer and can verify it yourself. Knowledge lookup is not the only use. For example "rephrase this text", or "give me ideas for X" just save you time and if the answer is wrong you can try again.
I don't see LLMs replacing traditional Search, but they still have plenty of other use cases.
and then it recommends a 3rd party library to solve your problem that doesn't exist
Apple is really slow on catching up on things, while they are good in introducing "new" things.
This is CPU-only RAM. It's way too low bandwidth to be used in high performance AI inference. The point with unified memory is that it's high bandwidth and available to the CPU, NPU, and GPU at the same time.
It should beat the base M2. It's the highest-end AMD laptop chip compared to Apple's lowest end. However, I have doubts that the Ryzen 7940 has a faster GPU than the base M2.
And of course, the M2 is significantly better in perf/watt.
Benchmarks seem to put the 7940 ahead of even the M2 Pro: https://www.cpubenchmark.net/compare/5454vs5189/AMD-Ryzen-9-...
Performance per watt seems remarkably similar.
Unified memory that is also high bandwidth.
>Benchmarks seem to put the 7940 ahead of even the M2 Pro:
Use Geekbench 6. It's closest to SPEC and optimizes well for both x86 and ARM.
Performance per watt is significantly in favor of the M2.
I do think Apple is smart in using memory the way they are. It's the future. But AMD routinely posts higher performance numbers for CPU and GPU because their large caches boast similarly high memory bandwidth figures and AMD's cores have significantly higher IPC. I also think Apple's trick of extending the ARM ISA to allow for x86 style memory alignment is slick for fast emulation. Credit where it's due. Apple's caches are also very fast, just not as large as AMD's. In short, memory bandwidth is just one measurement of the bandwidth along the entire disk bandwidth <-> disk cache bandwidth <-> PCIe bandwidth <-> main memory bandwidth <-> L4 / L3 / L2 / L1 cache bandwidth hierarchy, and details about cache coherency, associativity, eviction strategies, branch prediction, and other sundries matter along the way.
My 8 core Ryzen 5800x3D certainly compiles code a LOT faster than my 8 core M1 Mac despite the difference in main memory bandwidth.
More memory is almost always going to mean that memory is slower, and Ryzen can handle more memory. Ryzen's architected with large caches and prefetching and all kinds of other goodies with that in mind.
> Use Geekbench 6. It's closest to SPEC and optimizes well for both x86 and ARM.
Geekbench is a synthetic benchmark. I'd prefer real workloads like compiling code or tokens / s inference.
This is most certainly wrong. Apple Silicon has way higher IPC. AMD chips clock much higher, which uses much more energy. Hence, Apple Silicon is much more efficient.
>My 8 core Ryzen 5800x3D certainly compiles code a LOT faster than my 8 core M1 Mac despite the difference in main memory bandwidth.
Code compilation is not bound by memory bandwidth. But AI training and inference is. Hence, we're talking about it.
>Geekbench is a synthetic benchmark. I'd prefer real workloads like compiling code or tokens / s inference.
You send me a link to a synthetic benchmark. All I did was link you to a better and more accurate one.
Also, Geekbench performs real world workloads. You can easily look at individual scores for each test - besides the main score.
https://openbenchmarking.org/vs/Processor/Apple+M2,AMD+Ryzen...
Sure seems like the M2 is very far behind the 7700 in performance in all areas. Farther than can be explained by the difference in Mhz or package power.
Don't think Intel, AMD, or phone/laptop/desktop ARMs (from anyone but apple) have any near term plans for great memory bandwidth.
Sad. Obviously AMD can do it, they ship a much nicer memory interface on the PS5 and XboxX, but don't ship anything over 128 bit until you move up to $$$$ workstations with the threadripper (256 bit), threadripper pro (512 bit) or Epyc/Genoa (768 bit) wide memory. Even the Epyc/Genoa socket @ 460GB/sec doesn't match the M2 ultra (800GB/sec), despite the bare CPU being more expensive than most of the Mac studios sold.
M2 Max lands just around the 3060 mobile performance profile in this instance, it would be curious to see how the tokens/s reflect that.
Second, M2 Max might have a lot more RAM available than a 4090 because of its unified memory architecture. It can have up to 96GB vs 24GB for the 4090. So we need to look at different RAM configurations.
By Apple.
> unified memory architecture
Not really a bottleneck when layering over PCI exists. You're only constrained by PCI bandwidth and memory speeds, neither of which are really slow enough to meaningfully impact AI inferencing performance. Honestly, disk speeds are the #1 AI bottleneck I've seen on older systems.
> It can have up to 96GB vs 24GB for the 4090
Good, at the M2 Max's price point I could almost afford 4x 3090s anyways. I'd only need one to beat it in inferencing performance though.
However I suspect the unified memory architecture of the M2 is a big benefit here and this M2 Max system is either a 64 or 96GB model?
Even the very highest end Nvidia dGPU powered laptops won't have anywhere near that amount of GPU memory available and I don't know what the performance impact is with using system memory vs GPU memory on such systems?
> Watching llama.cpp do 40 tok/s inference of the 7B model on my M2 Max, with 0% CPU usage, and using all 38 GPU cores.
> Getting 24 tok/s with the 13B model
> And 5 tok/s with 65B
Which is still incredible. I recently tested llama.cpp on a dual xeon server with 24 cores and 70gb RAM and on the 65B model it took like 15 seconds for one token.
GPU seems the way to go here
here's the linked github fwiw
Btw, it's time to measure performance of LLMs in watts.