This is not a valid argument. TPS is essentially QoS and can be adjusted; more GPUs allocated will result in higher speed.
(Confirmation is faster than prediction.)
Many models architectures are specifically designed to make this efficient.
---
Separately, your statement is only true for the same gen hardware, interconnects, and quantization.
Speculative decoding is just running more hardware to get a faster prediction. Essentially, setting more money on fire if you're being billed per token.