I'm just thinking out loud. Nothing in this post is authoritative.
Theoretically, the time consumed to inference a single token with part of the model stored in flash should be equal to the time to inference that token if the whole model was in RAM, plus the time required to load the part of the model stored in flash memory.
I assume we do not need to write back to flash, but I'm not an LLM expert so I could be wrong.
I assume we have many (more than 10) layers so we can leave a fairly small amount of our RAM available to load one layer after another. Most nontrivial LLMs have many dozens of layers, so this seems plausible.
If we are not bottlenecking on our RAM during inferencing then we might be able to load the next layer from flash into RAM through a DMA transfer while inferencing our current layer. I don't think that would work on single-processor systems due to us always bottlenecking on RAM.
Maybe a dual-processor system could load one layer into RAM on one processor while inferencing on the previous layer on the other processor, and thus run really big LLMs in a small amount of RAM?
I'm sitting next to a pile of parts to build a new LLM AI machine. (z840, dual processor) and I look forward to playing with this stuff.