I'm not convinced about that 10 cheaper a year.
Larger models need more memory. I'm willing to bet that most of the tier 1 providers rely on multi-GPU models to serve traffic.
None of that is cheap, 8x GPU nodes that serve less than 20 queries a second are exceedingly expensive to run.