I am using the unsloth 4bit quants for both, with quantization aware training for gemma. I haven't tried other quants with these models. I also use a q4 quantized KV cache.
The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.