This wasn't a waste of time for me at all. On llama.cpp I could only run the IQ3_XXS quant at 21 t/s on my R9700 + system RAM, and on Strata I can run it at 60 t/s. Also... QFN at IQ3_XXS is giving better results for me than 27B at Q6 fwiw.
No comments yet.