Hard to fully predict the future, of course, but changes are the answer is "no".
Assuming you are comfortable with the current model performance rather than always trying to have the most performant model...
1. Lots of researchers are looking at ways to reduce model size, inference time, etc... We see smaller models outperforming older benchmarks/achievements. Look into model distillation to see how this is done for specific benchmarks, or new approaches like Mixture of Experts (MOE) that reduce compute burden but get similar results as older models.
2. If your concern is cost of running a model, then GPU tech is also getting faster/better so even if compute requirements stay the same, the user will get an answer faster + will be cheaper to run due to the new hardware.
I hope that helps!