https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...
https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...
Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).
DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.
GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.
The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.
It's the first small local model I've felt like I can do real work with.
unsloth/Qwen3.8-27B-GGUF UD-Q3_K_XL DSH (pi)
Any tips?
we went from 62% completion to 92% using a claude code harness
I use the Unsloth UD_Q2_K_XL GGUF with default parameters, along with that custom template linked elsewhere in the thread, and no K/V cache quantization.
Try the same prompt with a larger quant (even if it runs very slowly because the model no longer fits in VRAM) & see if Qwen does better - if so, there’s your answer.
My use-case is coding, currently working on a project with a Rust backend and TS/React/Vite frontend, with probably tens of thousands of lines of code total (including tests).
Where Opus has seen it before and knows how to do it, Qwen knows how to work it out. It turns out it's surprisingly capable at working things out. The obvious drawback is that it takes tokens and time.
All the same, getting to run something this capable locally is momentous, and suggests to me that streaming tokens from colossal data centers might not be the long term path forward.