I believe lambda.cpp has been designed for at least an M1 - no idea if there are options for running LLaMA on older hardware.
If it used the GPU/ANE and was a true large language model then it would only work on M1 systems because they're unified memory (which nothing except an A100 can match.)
LLaMA's GPT-3 175B level model, LLaMA-13B, only requires 8GB of VRAM (and no RAM) to run using pre-quantized 4bit weights. So it's hardly a job for an A100.
Even the largest model, LLaMA-65B, is only 30GB and inference can be split between two graphics cards with almost no effect on performance.