I'm not saying everyone has to run local LLMs, because the APIs are in a race to the bottom, and my $10 of OpenRouter credits I bought months ago is down to $8.94 because most models give you MILLIONS of tokens for a US Quarter.
This is tunnel vision. The percentage of people who could afford the hardware you could at the time you back it so vanishingly small. I do not know a single non-tech person who has multiple graphics cards in a single computer.
My personal expectation is closer to 5 years than 10, which is why I wouldn't touch Anthropic or OpenAI stock with a ten-foot pole, personally, no matter how high their theoretical valuation is. Because their business model is doomed in the long run.
Which means it's inevitable that eventually, even the consumer game market will be buying GPUs with 32 or 64 GB of RAM. And there are decent models that will run at that size. Even the "normies playing games" market, as you call it, will end up with the capacity to run local models. It'll take a few more years than it would have if the data-center companies weren't trying to buy up all the GPUs, but it's not like gamers are going to stop wanting to play games. So in the long run, Anthropic et al are still going to have to figure out how to deal with competition from local models that run on your gaming video card. Which won't ever be at parity with the models that take terabytes of VRAM to run, but are very rapidly approaching "good enough for what most people want to do".