This wasn't a waste of time for me at all. On llama.cpp I could only run the IQ3_XXS quant at 21 t/s on my R9700 + system RAM, and on Strata I can run it at 60 t/s. Also... QFN at IQ3_XXS is giving better results for me than 27B at Q6 fwiw.
Quants vary by model. DS4 is very credible at a 2-bit quant. Not sure about Qwen 3.8 Flash Next; I run it at a 4-bit quant and it's too slow, so I'm trying out DwarfStar today to see if that improves things.