Qwen3.8 27B beats all medium models (40B–150B). It has the same score as DeepSeek V4 Flash 0731, which ranks #5 in large model category (> 150B).
Sources:
- https://artificialanalysis.ai/models/open-source/small
Qwen3.8 27B beats all medium models (40B–150B). It has the same score as DeepSeek V4 Flash 0731, which ranks #5 in large model category (> 150B).
Sources:
- https://artificialanalysis.ai/models/open-source/small
https://simonwillison.net/2026/Aug/16/qwen-38-27b/
It seems like the token usage is 2.3x GPT Luna Max and almost 2x Kimi K3!
I'm curious if they can make up for this with insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!)
https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B
And then a Bonsai ternary on top of that model.
...or did Prism do something special with their "bonsai" releases? I didn't notice anything like QAT being mentioned.
On the other hand, there are some of us who are stuck with hardware that has plenty of compute, but limited (V)RAM. The new 27B is just perfect for that.
No? Gemma 31B and Qwen 27B are about the same speed. Gemma 26B-A4B and Qwen 35B-A3B are about the same speed.
Output TPS in vllm for instance:
- Gemma4 26B-A4B: 200-300TPS
- Qwen3.6 35B-A3B: 120-180TPS
- Gemma4 31B: 80-120TPS
- Qwen3.6 27B: 60-80TPS
This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.
| model | size | test | t/s |
| ------------------- | ------- | ------ | ---- |
| gemma4 31B Q4_0 | 16.1 GB | pp2048 | 1248 |
| gemma4 31B Q4_0 | 16.1 GB | tg512 | 40 |
| qwen35 27B Q4_K | 15.9 GB | pp2048 | 1248 |
| qwen35 27B Q4_K | 15.9 GB | tg512 | 39 |
| gemma4 26B.A4B Q4_0 | 13.3 GB | pp2048 | 4304 |
| gemma4 26B.A4B Q4_0 | 13.3 GB | tg512 | 160 |
| qwen35 35B.A3B Q3_K | 15.7 GB | pp2048 | 3329 |
| qwen35 35B.A3B Q3_K | 15.7 GB | tg512 | 144 |
> with their respective speculative decoding methodsYou're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower.
One note I had between the two is that gemma has a much higher prefix cache hit rate in general.
I gave general numbers of what I'm getting above, the performance ratios seemed similar regardless of setup (eg. getting a AWQ-in4 quant on a single GPU vs PP without speculative decoding vs TP with speculative decoding).
Overall single GPU is fastest, and TP+speculative decoding is still faster than PP, but for fp8 models you need dual GPUs whether you want it or not.
./llama.cpp/llama-server \
-hf unsloth/Qwen3.8-27B-GGUF:UD-Q8_K_XL \
--webui-mcp-proxy \
--no-mmproj \
--parallel 1 \
--kv-unified \
--flash-attn on \
--fit off \
--split-mode tensor \
-ngl 999 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
-ub 256 \
--no-context-shift \
--host 0.0.0.0 \
--tools all \
--jinja \
--ctx-size 262144 \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--reasoning on \
--chat-template-kwargs '{"reasoning_effort":"medium"}' \
--reasoning-preserve \
--temp 1.0 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--presence-penalty 0.0 \
--repeat-penalty 1.0
Use claude/codex/whatever with /goal to optimize params for you.IMHO draft model support on dense models is great alternative to MoE on GPUs (high bandwidth, less memory) – more intelligence, speed somewhere mid way there which is usually sufficient.
Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out:
```
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
vllm serve Qwen/Qwen3.6-27B-FP8 \
--dtype auto \
--kv-cache-dtype fp8 \
--enable-chunked-prefill \
--enable-prefix-caching \
--trust-remote-code \
--enable-auto-tool-choice \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":3}' \
--default-chat-template-kwargs '{
"enable_thinking": true,
"reasoning_effort":"medium"
}' \
--tensor-parallel-size 2 \
--max-model-len 250000 \
--gpu-memory-utilization 0.9 \
--max-num-batched 12000 \
--max-num-seqs 24
```I took the liberty of adding your reasoning effort chat template to my setup. You can play around with the last few parameters. In generall VLLM will be better in higher concurrency scenarios, so if you only use it for a personal vibe coding assistant and less as a general home model for task execution llama.cpp may be better.
it's good to play with harness setup where you fan out multiple concurrent branches that share non trivial amount of prefix then reduce their output/summary back into main agent.
ie. instead of serially reading further skills/relevant source files for planning/thinking, you can branch and read them in parallel reusing prefix / or use to to approach request from different angles in parallel - to map-reduce result onto main context of what's actually relevant. branching subagents has benefits of not polluting main context, shared prefix prefill is close to free on a cache hit and with concurrent decoding/continuous batching you can utilize gpu well to get good speedups.
ie. what's relevant is number of active concurrent sequences (and their shape, ie. shared prefix), not so much number of users.
i'm not sure with llama.cpp vs vllm regarding concurrency – llama server has multiple server slots, continuous/dynamic batching enabled by default, prompt caching (also on by default), ram prompt cache, context checkpoints, unified kv buffer across sequences etc. so shouldn't be bad, i guess would be good to actually benchmark. personally i'm happy with llama.cpp.
It's a dense model so it will use all of its parameters per token. 37B active parameters isn't tiny at all, it's almost what Deepseek R1 had, and it's 2/3 of what Kimi k3 uses, so it's not going to be “insanely high” tps: it's going to be three times slower than Deepseek Flash (Prefil speed is going to be quite high though, but not token generation).
The fact that we have a GLM 5.2-class model that can run on two 3090's comfortably at Q8 is absolutely insane. It wasn't long ago that GLM 5.2 was considered amazing for open weight models.
xhigh -> "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer."
medium -> no mention of effort (sentence omitted)
low -> "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration."
in my testing this doesn't seem to produce exactly deterministic thinking levels, because it's just a system prompt nudge. i had instances where medium thought longer than xhigh
You can also apply fixed token budgets for the reasoning blocks, though it will decrease quality in some cases.
But I think the next step will be even more thinking on smaller models. Maybe fine-tuned and we get really crazy stuff