Performance of llama.cpp on Apple Silicon A-series
github.com
github.com
Given how big these models are (and the steep cost for GPUs to load them), I had been thinking that most people would interact with them via some hosted API (like what openAI is offering) or via some product like Bard or Copilot which offload inference to some big cloud datacenter.
But given how well some of these models perform on the CPU when quantized down to 4, 6, or 8 bits, I'm starting to think that there will be quite a few interesting applications for fully local inference on relatively modest hardware
That being said your less technical users will probably still use them remotely, and the very largest models will also probably be hosted just because of the capital cost. It doesn’t make sense to spring for a super high end rig to run a massive model locally unless you are using it very very heavily or are a serious homelab enthusiast willing to shell out some bucks.
Of course there are wildcards here like breakthroughs in model efficiency, compression, or distributed execution and training. Any of that could change the game. I get the sense that we really don’t know how far we are from optimal efficiency right now. There may be far more efficient architectures or ways of compressing these things.
Part of me envisions a Coral-like device, but the kicker/difficulty/cost problem is always going to be fast, big memory. I think that puts a bit of a price floor on some of it. Can't wait to be wrong though!
> how about a co-op neighborhood LLM rack?
It'll work about as well as those "wire your neighborhood for Internet as a collective" movements of the 1990s.In fact, you'll need those to deal with latency issues, if my experience with consumer ISP quality is any indication.
1. https://store.ui.com/us/en/pro/category/all-wireless/product...
Local inference is quite inevitable.
You're selling a game once rather than selling a subscription, so you're locked in to paying for cloud-based LLM services indefinitely, unless you want to shut the game down after X years, which really annoys people. Also, you're incentivised to be stingy with the LLM-based features because the more you use it the more it costs you, but by running it locally you can offload that cost to the consumer.
As an experiment, I've been running llama.cpp on an old 2012 AMD Bulldozer system, which most people consider to be AMD's equivalent of Intel's Pentium 4, with 64 gigs of memory, and with newer models it's surprisingly usable, if not entirely practical. It's much more usable, in my opinion, than spending energy trying to get everything to fit in to more modest GPUs' smaller amounts of VRAM.
It certainly shows that people shouldn't be dissuaded from playing around just because they have an older GPU and/or a GPU without much VRAM.
Is that separately comparing the time it takes to preprocess the input prompt (prompt_length / pp_token_rate = time_to_first_token) and then the token generation rate is the time for each successive token?
I also see something about bs batch size. Is batching relevant for a locally run model? (Usually you only have one prompt at a time, right?)
it makes sense to benchmark them infependently since prompt processing is done in parralel for each token and is compute bound and token generation is sequential and bound by memory banwidth
Prompt processing doesn't actually need to compute logits for each new token, just cache the KV values and so is much faster than actual inference.
Generation is the response it sends back
But "AI" is still very research focused (training) and not application focused (inference). Llama.cpp et al are filling in this gap.
8x7B nowadays
As long as metal is used on an iphone I could see it worked well too. I use quantize 5 on my laptop but quantize 4 seems very practical
$ ollama run mixtral
Apple’s chip architecture has a variety of co-processors alongside the CPU that make it the most perfect for transformers, in a convenient laptop form factor with fanless heat dissipation and low energy footprint
There should be 10-20 input and output that is tested for correctness or something in addition to t/s as a reference