I would think so, I have it quantized at 8 bit (q8) and that ticks in at ~47GB.
Q4 should be well below the 32GB (2x 16GB).
I am wondering if the same implementation could be done for RAM to CPU, too. Those are said to be data transfer limited, so minimizing the RAM to CPU cache transfer should help there, too?