I have a small laptop.
If you have more disks available, you could really do some testing.
When you have some benchmarks, submit a pull request or issue so we can maybe work on them.
We are really happy for contribute!
I think another route might be looking at holding an even larger chunk of model weights in ram, and taking advantage of RAM<->GPU bandwidth, perhaps using a PCIe 5 GPU. This was my first thought since I have dedicated GPU.
If you are using Laptop, you're looking at shared memory between the iGPU and CPU. I've also tried that route, but I have always been skeptical of killing flash with too many reads, it essentially uses SSD like it's a consumable item.
I'm going to benchmark this right now with what I have and I'll get back to you on github.