This is neat, but what I really want to see is someone running it on 8x 3090/4090/5090 and what is the most practical configuration for that.
You can rent a single H200 for 3$/hour.
Still, 8x3090 gives you ~2.25 bits per weight, which is not a healthy quantization. Doing bifurcation to get up to 16x3090 would be necessary for lightning fast inference with 4bit quants.
At that point though it becomes very hard to build a system due to PCIE lanes, signal integrity, the volume of space you require, the heat generated, and the power requirements.
This is the advantage of moving up to Quadro cards, half the power for 2-4x the VRAM (top end Blackwell Quadro expected to be 96GB).
Benchmarks: https://github.com/ggerganov/llama.cpp/issues/11474#issuecom...