Efficient LLM inference solution on Intel GPU
arxiv.org
arxiv.org
Is this paper the same work as goes in to Intel's BigDL-LLM code? That's been out for a few months now but I haven't seen it in use yet. https://medium.com/intel-tech/bigdl-llm-easily-optimize-your...
Neither have I, but this is interesting and I am bookmarking it, thanks.
The Gen AI space is littered with very interesting high performance runtimes that precisely no one uses. There are too many for integrators to even keep up with, much less integrate!
The hardware seems stuck in past decade and the process woes don't help either - but the software should be ready if they ever dig themselves out of the hardware hole.
And then I may end up fixing Intel bugs in ML projects we use, and deploying them on Intel cloud hardware... which is hopefully more available as well.
I’ve been working with their research teams on this and credit due indeed.
Some context: https://github.com/ggerganov/llama.cpp/issues/2555#issuecomm...
Given the article about GPUs, I think I was thinking more along the lines of a desktop with a consumer GPU.
The CPU/GPU speed of the Air is the same as the MacBook Pro base model though. The only difference is the lack of active cooling, which for large workloads can result in performance degradation. Chat apps are intrinsically interactive though, only using bursts of GPU when it is performing inference. Shouldn't be an issue.
The expensive PRO/MAX variants of the MacBook Pro would be 2x - 4x faster. But the plain M2 in the MacBook Air is sufficient to get real-time speeds on sizable models.
When the M3 comes out for the MacBook Air, discounted M2 machines, new or refurbished, should be an ideal entry-level machine for local inference.
At the point of purchase of the lowest cost configuration with 24GB Unified Memory, you've already paid the an equivalent of over 2200 hours of GPU compute time on an RTX 4090 24GB, with a performance that exceeds the MacBook by around 1200% (it/s).
If you buy that MacBook for AI, you would have to run continuous generative inference on it for over a decade to match the return of just having used runpod instead. Even Apple doesn't use Macs to do AI - they use GCP TPUs. Buy the Mac if you like it, by all means, but be realistic.
LLM performance on Apple Silicon is decent for small models but does not scale. If any Mac architecture were cost effective at any scale for AI, we would be putting them in data centers, and we're not.
It’s an entry-level Mac, and it is capable of doing inference jobs. If you have around $1k available to spend on a daily driver laptop, but you also want to do some AI inference experiments, the MacBook Air is what I’d recommend.
The AMD Framework laptop might be a good alternative. But I don’t know how well its integrated GPU is supported by llama.cpp.
Its quite smart, and fast.
The whole rig cost me $2.1K, but it could have easily been $1.5K without splurging on certain parts like I did. And its 10 Liters, small enough to move around.
https://www.reddit.com/r/sffpc/comments/18a7mal/ducted_3090_...
Screenshot: https://github.com/Const-me/Cgml/blob/master/Mistral/Mistral...