thanks for posting your setup! I think it's smart to set the reasoning effort default to something saner in the base config.
Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out:
```
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
vllm serve Qwen/Qwen3.6-27B-FP8 \
--dtype auto \
--kv-cache-dtype fp8 \
--enable-chunked-prefill \
--enable-prefix-caching \
--trust-remote-code \
--enable-auto-tool-choice \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":3}' \
--default-chat-template-kwargs '{
"enable_thinking": true,
"reasoning_effort":"medium"
}' \
--tensor-parallel-size 2 \
--max-model-len 250000 \
--gpu-memory-utilization 0.9 \
--max-num-batched 12000 \
--max-num-seqs 24
```
I took the liberty of adding your reasoning effort chat template to my setup. You can play around with the last few parameters. In generall VLLM will be better in higher concurrency scenarios, so if you only use it for a personal vibe coding assistant and less as a general home model for task execution llama.cpp may be better.