I only trust those users genuine personal tests
I only trust those users genuine personal tests
https://www.youtube.com/@lukesdevlab
I don't know if that is what you are looking for or not and as always your experiences may be different.
However, watching tests of heavily quantized models that weren't designed for it (non-QAT) is frustrating. There's no way to tell if the actual model fails because it's dumb or if the lobotomy made it that way.
running an untouched, vanilla 4-bit version (Q4_0) I baked myself today (benched it against Q4_K_M (16gb) and IQ3_M (12gb), Q4_0 (15gb) is king)...
this model--
1: over 60% faster than qwen 3.6 version of the same dense 27b model, same engine setup (don't ask how, im not sure either)
2: has better reasoning quality, less "loopy" with its thinking patterns.. most certainly the smartest model on my roster currently
3: has the longest task horizon ive ever experienced (locally or otherwise)...I sent it a bunch of compressed ideas for an app, it sent me back the largest python app ive ever seen in one single ai response pass (80kb text file)
thanks qwen!! hoping to see the full model range get released...