- memory compression algorithms
- alternative LLM architectures that don't rely on memory or GPUs
- compatibility hardware (like DDR3 to DDR4 boards)
- distributed computing improvements, both at local GPU and networking levels (SLI for AI)
- GPU hacks to add more memory or support older architectures
I'm personally looking forward to the new LLM architectures that don't require as much compute, e.g. DLLMs, which can be good enough for CPU usage but lack the accuracy of frontier models currently.
When this happens the bottom will fall out of the GPU and memory markets, putting a glut of cheap hardware out there.
Doom mongering like this never seems to include these as viable future alternatives, which is standard market adjustments, I wonder who the doom narrative helps? :)