How’s it looking for a six year old MacBook?
Not there yet?
Does this still use the gpu?
Not there yet?
Does this still use the gpu?
If it used the GPU/ANE and was a true large language model then it would only work on M1 systems because they're unified memory (which nothing except an A100 can match.)
LLaMA's GPT-3 175B level model, LLaMA-13B, only requires 8GB of VRAM (and no RAM) to run using pre-quantized 4bit weights. So it's hardly a job for an A100.
Even the largest model, LLaMA-65B, is only 30GB and inference can be split between two graphics cards with almost no effect on performance.