Are you willing to share any lessons learned, etc. that I could make use of? We are evaluating paying for a SOTA sub or trying this, and the talk about Qwen3.6-27B makes me want to try deploying this machine.
Are you willing to share any lessons learned, etc. that I could make use of? We are evaluating paying for a SOTA sub or trying this, and the talk about Qwen3.6-27B makes me want to try deploying this machine.
It's not even a real comparison if they are actually using them for coding.
If you are deploying always running agents (e.g. monitoring logs and services) then sure - a QWEN local server is a good choice. But for coding the cost in productivity of using a lower performing model is way too high.
But in term of actually running a dev team - you are free to use QWEN or another quantized local model that can run on an RTX 5090 for coding if it makes you feel more independence. However you would struggle and spend many many more hours achieving the same thing, with a lot more debugging time, long delays before it's done, and many more prompts.
It's just not the right approach. I use QWEN and other local models all the time, but for more clearly defined monitoring and classification tasks.
OTOH, as we see, the larger models demonstrate diminishing returns, smaller models demonstrate improvements, and hardware does not show any signs of becoming cheaper, so holding on existing decent GPUs may, too, be a winning strategy in longer term.
For continues all day work you definitely need a higher tier sub level.
I'm actually looking into deploying a GPU at my company because we can not give out our code. Qwen 3.6 looks good
For some projects the giving out the code part might be ok (i use Codex there too) but for the core app at the company I'm working at there is currently a strict no-AI policy. A local GPU solves this.