Qwen3.8-Next, thanks to its new architecture, is quite fast even if part of it is streaming from disk
Totally. Any MoE model can have experts swapped in and out from disk or system ram. I only framed it this way because the question was about the model fitting in vram.