This is where inference speed starts to matter. H100 might be cheaper per inference than Groq but cutting down the wait time from 1 minute to 10 seconds could be a big deal.
The hardware (or OS) is useless if it doesn't give good results, even if it theoretically could.