GPUs in the future will get more powerful. Especially if they are designed with LLMs in mind.
I assume there will be advances that basically make the context window practically infinite. That ought to make the lower parameter count models much more powerful on its own. I also assume they will become more efficient to run through sparsity/pruning, though I'm a total novice on this topic.
I wonder how many generations of hardware are we talking? One? Two? Or three? It feels like it's potentially within that range.