The cutting edge, max size models will likely stay in the GPU space for a long time.
But these models are not needed for most general requests.
With a fine tuned 30B quantisized model you can serve a large portion of requests with around 32GB of RAM.
Free users will likely only get these kinds of models.
At some point we will get these models in hardware and the cost per token will be minimal.