hey, I’m the author. That box has 384gb, but loading the model “only” uses about 80gb.
For CPU inference on old hardware I don't think q4 offers any benefit over q8 since the AVX unit doesn't support such small floats. I don't even think AVX supports 4-bit int math. IIRC AVX2 does.