>which nothing except an A100 can match
LLaMA's GPT-3 175B level model, LLaMA-13B, only requires 8GB of VRAM (and no RAM) to run using pre-quantized 4bit weights. So it's hardly a job for an A100.
Even the largest model, LLaMA-65B, is only 30GB and inference can be split between two graphics cards with almost no effect on performance.