Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules for higher perf. Are you using TP?
I gave general numbers of what I'm getting above, the performance ratios seemed similar regardless of setup (eg. getting a AWQ-in4 quant on a single GPU vs PP without speculative decoding vs TP with speculative decoding).
Overall single GPU is fastest, and TP+speculative decoding is still faster than PP, but for fp8 models you need dual GPUs whether you want it or not.