V100S 32GB, I have had Claude optimizing it for about a week and it is already at around 900 t/s prefill, 90-100 t/s output in Pi on coding tasks. There is also a Ninfer fork for the v100 but it requires a custom format. I am working on upstream Unsloth with GGUF 4-bit quant.
(I also have flash next running even faster on this machine, something a single 5090 can do, with expert cache/pinning, but not quite as fast) :)