The whole industry is now pushing through memory.
In 5 years you have either some type of explosion which willjust make all the hardware from today affordable or you have such an AI explosion, that the today hardware is written off and not efficient enough anymore that you can buy it for cheap.
In parallel, its clear that we need more memory.
In parallel models in hardware will become a thing on mass market.
In parallel everything gets more efficient. The 30B parameter model will be for sure more intelligent in 5 years than it is today.
It's in large part a problem of bandwidth, too. Mostly really. HBM memory can do up to 3TB/s vs DDR5 like 250GB/S. The latter is just too slow to process 2.5B parameter models, it simply can't move the values back and forth fast enough. It would drag to a crawl. Much smaller dense models at that speed on the NVIDIA Spark can't do more than 15tok/sec.
Real serving systems for these models involve large numbers of parallel GPUs with massive memory bandwidth, hooked up via NVlink.
It will take a long time for that level of tech to get down to consumer level.
(An ideal computing architecture built for LLMs would in fact offer some way of colocating computation with memory. If you can put matmul etc right in the DRAM and avoid going back and forth over the bus...)