"Yes, I'm using AI to speed up development. I'll submit that PR in two weeks time!"
Not only in big desktops, but even in most recent mini-PCs, it is possible to read concurrently from one PCIe 5.0 SSDs and one PCIe 4.0 SSD, at a total sustained reading throughput of around 20 GB/s.
With an optimized inference implementation, it should be possible to overlap completely the computations with streaming weights from SSDs.
This should improve the inference speed to around at least 1 token per second, on a cheap computer, under $2000 even at the current super-inflated prices.
There are enough tasks where this would be useful. Obviously one should use for most tasks a fast small LLM and use the big one only when this actually saves time.