Qwen3.6-Plus: Towards real world agents
qwen.ai
qwen.ai
Comparing to Opus 4.5 instead of the current 4.6 and other last-gen models is clearly an attempt to deceive, which isn’t winning them any points either.
I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. I don’t know how successful they’ll be in the race to the bottom in this market niche, though. Most users of cheap API tokens are not loyal to any brand and will change providers overnight each time someone releases a slightly better model.
There isn't, pretty much everyone wants the best of the best.
There are a lot of data science problems that benefit from running the dataset through an LLM, which becomes bottlenecked on per-token costs. For these you take a sample subset and run it against multiple providers and then do a cost versus accuracy tradeoff.
The market for API tokens is not just people using OpenCode and similar tools.
For direct user interaction or coding problems, perhaps. But as API calls get cheaper, it becomes more realistic to use them for completely automated workflows against data-sets, or as sub-agents called from expensive SOTA models.
For example, in Claude, using Opus as an orchestrator to call Sonnet sub-agents, is a popular usage "hack." That only gets more powerful, as the Sonnet equivalent model gets cheaper. Now you can spawn entire teams of small specialized sub-agents with small context windows but limited scope.
Seems like a huge waste of money and electricity for processes that can be implemented as a traditional deterministic program. One would hope that tools would identify recurrent jobs that can be turned into simple scripts.
For example: "Here our dataset that contains customer feedback comment fields; look through them, draw out themes, associations, and look for trends." Solving that with a deterministic program isn't a trivial problem, and it is likely cheaper solved via LLM.
I did create my own MCP with custom agents that combine several tools into a single one. For example, all WebSearch, WebFetch, Context7 exposed as a single "web research" tool, backed by the cheapest model that passes evaluation. The same for a codebase research
Use it with both Claude and Opencode saves a lot of time and tokens.
One data point that might help with the research tool: not all APIs work equally well when called by agents. We scored 387 APIs on agent-friendliness and 54% fail the bar. The main gaps are no CLI tool (66%) and no machine-readable pricing (72%). If your research tool is helping agents pick APIs to integrate, the scores at clirank.dev/api/apis?sort=score could save the expensive model from wasting tokens on APIs that'll fail headless.
There are many simpler tasks that would work fine with a simpler, local model.
At least from my experience and friends of mine, we use OpenRouter for cases where we want to use smaller LLMs like Qwen, but when I've used ChatGPT and Claude, I use those APIs directly.
I think it's pretty disingenuous to call your SaaS little when it is projected to spend at least 5 million USD just on tokens and this is a low end estimate.
And I pay way less than $1 per input token, especially when caching is taken into account.
EDIT: they updated it in the last day or two, now it says 70T, so I’m a little below 0.1% now. But seriously, the point stands, 70T tokens a month just isn’t that much in the global scheme. The big labs are pushing quadrillions each.
Coding is a rung on the ladder of model capability. Frontier models will grow to take on more capabilities, while smaller more focused models start becoming the economical choice for coding
products entirely disappearing or significantly changing will be more and more common in the llm arena as things move forward towards companies shutting down, bubbles deflating, brand priorities drastically reshifting, etc...
i think, we're at or at least close to a time to really put some thought into which pieces of your flow could be done entirely with an open/local model and be honest with ourselves on which pieces of our flow truly needs sota or closed models that may entirely disappear or change. in the long run, putting a little bit of thought into this now will save a lot of headache later.
But with LLMs, how do you know switching from one to another won’t change some behavior your system was implicitly relying on?
I have no affiliation with DeepInfra. I use them, because they host open-source models that are good.
GCP does server other non-Google models, but I'm not sure what they have other than Anthropic models. I don't think Haiku is a great model though.
The price is a concern too of course. But privacy is a bigger one for me. I absolutely don't trust any of their promises not to use data for training purposes.
I would say that for a significant part of the current market open-source models are good enough to fill a part of it.
Right, they state that they'll release "smaller" variants openly at some point, with few details as to what that means. Will there be a ~300B variant as with Qwen 3.5? The blog post doesn't say.
Now they show their true colors. They want to train models on our engineering to replace us, while simultaneously giving nothing back? No thanks. I'd rather fund the shitty US hyperscalers. At least that leads to jobs here.
If there's a company willing develop and foster large scale weights in the open, I'll adopt their tooling 100%. It doesn't matter if they're a year behind. Just do it open and build an entire ecosystem on top of it.
The re-AOLization of the internet into thin clients is bullshit, and all it takes is one player to buck the rules to topple the whole house of cards.
That's a very reasonable stance. It doesn't change the fact that we do have plenty of local models (up to and including Qwen 3.5) that are still quite useful.
I'd prefer them to be open weight, but I'd love to sub a decent competitive coding plan from a European or Chinese provider. Right now they're not quite there. If closing it and charging for it brings them closer to competitive, that's ok.
If the US tech and AI industry long term wants customers and a broad market outside of their own domestic base, they need to reconsider who they are bending the knee to, and how they are defining their policies in relation to the Trump administration.
Bring on the Chinese competition.
Bring on the Chinese, fuck the Americans.
Just without the bombs part.
Seriously, even if you manage to elect someone capable ever again, the US can't be trusted to not elect someone worse than a toddler again in 4 years. In fact, you can't even be trusted to not elect someone even worse than your current dictator (if you thought it couldn't get worse, Trump isn't the bottom of the barrel by far).
It will help them get a good flank on the USA such that even when that temporarily embarrassed country gets a leader you, and the rest of the world do like, it will be too late to do anything.
A perfect definition of cutting off your nose to spite your face laid bare for all to see.
And then, compared to China, the US acts overtly hostile: threatening us with war, starting a war in order to collapse energy supplies outside of the US. Opportunistic beyond even China, much more hostile.
Will the US even be a democracy in two years? Is it now?
Nah man, balancing between China and the US is the only thing a smaller country can do in order not to be crushed
We have an American neighbour actively funding and amplifying a formerly extremely fringe separatist movement in Alberta -- shades of the Donbas, North American edition --and a US "ambassador" who has the behaviour of a 4chan troll.
The bridge has been blown up. Americans might think they are a midterm election away from salvation, but we're on the whole not so naive.
And mostly not liking what we see. Encouraged by the No Kings protests, but unless that boils over into a hegemonic and stronger opposition, it still seems like there's a 40% population there that can't deal rationally with the world inside their own border, let alone outside.
Also... When Biden took over after Trump's first term most of the protectionist policies stayed and foreign policy didn't really budge (outside of support for Ukraine). I expect similar if (big if) the Democrats regain executive power.
codex handles longer sessions but the quality seems to decline and it tends to over engineer and lose focus. It will happily add slop on top of slop...which may pass immediate tests of "code works" but doesn't pass my criteria of "code as craft"
I'm using z.ai GLM with opencode. It's obvious when GLM loses its mind when the session gets too long.
I've been using AI to support programming for around 3 years now. The models have gotten amazing. However, unless there is a significant breakthrough I have determined that it's best for me to focus on short sessions.
I a) organize my work, b) improve my AGENTS.md, ensure source has appropriate comments to guide the models to the patterns and separation of concerns c) use shorter sessions d) review and test without AI. This approach means I still own my code. The AI is just an assistant.
With this approach GLM-5.1 is an excellent model. I never run out of token allotment on z.ai or codex plans. At this point, I only keep my OpenAI subscription as the ChatGPT desktop app is excellent at long web research tasks and I get codex with it.
Qwen is not the only Chinese lab, and the others have shown no change in their commitment to open source. Allegedly Qwen hasn't either if their recent statements are to be believed. They're just hoping to capture market share with *-claw customers before releasing an open weights version. We'll have to wait and see how before they decide to release that.
I wouldn't call this totally accurate, especially as of late. What's closer to the truth however is that there's lots of second-rate players in China doing open models, that will be getting a lot more attention from local AI proponents if the big names seriously slow down their AI releases. The local AI scene as a whole is quite healthy.
What exactly has changed? Alibaba just released a bunch of new models and have said 3.6 weights are coming soon. The others labs have shown no signs of slowing down their releases either. Whatever you're referring to is news to me and likely most others.
I'm USian myself, but I don't think the site should be very US-centric.
Only academic models will be true open source as companies can't legally afford to disclose learning inputs.
In regards to "They want to train models on our engineering to replace us". Some software engineers in China can run circles around some of the best teams in Silicon Valley. Days of U.S. hegemony are over. I recommend you make peace and make friends.
I can constantly jump from one provider to another, and to my local servers which are already able to run very useful models at reasonable hardware cost, and I intend to continue doing that for the foreseeable future
the one thing I'm not going to do is tying my tooling to one provider or another or getting overtly used to the specifics of a model outside of my control
more than the weights or the training, which of course are very important, the real battle right now is for establishing some dependency mechanism so that your users won't just flee en-masse as soon as you inevitably try to abuse your market dominance and lock-in mechanisms, as is customary in everything computing these days - note that i don't explicitly talk about raking up prices, that is just one of the most difficult methods as people are very sensitive to that, when you can sneakily sell user data, get government contracts from never-disclosed conditions, or even just incorporate your intelligence to ad networks in one way or another
This is how I view that the public can fund and eventually get free stuff, just like properly organized private highways end up with the state/society owning a new highway after the private entity that built it got the profits they required to make the project possible.
It's super not a publicity stunt, qwen 3.5 is the base of the best local models out there IMO.
There's smaller models all the way down too.
Like this should be _exactly_ what we want companies to release.
The naivety around this has been staggering quite frankly. All of a sudden, people thinking that meta etc are releasing free models because they believe in open access and distribution of knowledge. No, they just suck comparatively. There is nothing to sell. Using it to recruit and generate attention is the best play for them.
Apparently that wasn't actually the play here.
There's nothing really strange about not competing directly with the best, but rather showing whom you are as good as.
They posted charts with logos for Claude and others. You had to read the fine details to realize they weren’t comparing to the latest offerings from those companies. They were counting on you not noticing.
There’s zero reason to compare to old models unless you’re trying to mislead.
Sure they are not cheap to train. But if open weight models continue to be trained and continue to become available on cheaper hardware, how do dedicated AI companies protect their margins?
Rather than an increase in VRAM of consumer gpus, we are seeing a decrease, which is pretty optimal for OpenAI
The former gets stuck in ridiculous thought loops on the exact same tasks I’m testing. Fascinating really, I expected more for some reason.
I'd agree that 4.6 and 4.5 are different, but I don't think it's correct that 4.6 is just reduced and benchmaxxed. It genuinely solved problems for me that no other model has been able to.
I think I'd like to have seen the 4.6 benchmarks also included against Qwen.
So expect every now and then a open model burp from the trailing frontiers. Afterall, its all sunk cost so once you have it and no customers, theres zero reason not to spike your competition and try again or exit.
Chinese models are very competitive in that regard, you'll often look at 70-90% price reduction at the same quality.
In the exploration phase, yes. But once your setup settles down you likely want to stay on the same model for stable operation.
This field is going in a incredible pace, the providers release a new model every quarter or so. The amount of criticism is a bit overblown in my opinion. The benchmarks still look very good to me. I’ve used GLM-5 (latest is GLM-5.1) and Kimi K2.5, they are decent and gets the job done, so seeing how this model of Qwen performs compared to it is kinda impressive.
Also, why are so many pointing out the fact that this model is not open-weight as if this is their first time doing so. Qwen-3.5-plus, Qwen-3-Max is also closed source. This is not something new.
I think Qwen trying to catch up to the SOTA models is still healthy for us, the consumers. Sure, its sad news that this version is closed-weight, but I won’t downplay their progress.
Laziness? Lack of time? It's not like the latest generation of the SOTA models were released yesterday.
Now, is it mildly deceptive because all of the companies using incredibly confusing naming conventions for their models? Maybe!
I don't think any org doing this is necessarily being deceptive, so long as there's some reasonable basis for the chosen comparable(s).
For example, comparing a new iPhone to a prior Android phone might make sense if the install base is considerably large and Apple is targeting the cohort for user acquisition. (~"These benchmarks are not for you.")
The community will always run the numbers and get the clicks for the benchmarks not filled in by the 1st party. I noticed what appeared to be some movement from Apple in content they've produced to get ahead of this with recent product content.
Opus 4.5 is $25/m output tokens.
This is at most $6/m output tokens.
That's ~1/4 the price.
I used the https://modelstudio.alibabacloud.com/ API to generate that one, which required signing up for an account and attaching PayPal billing - but it looks like OpenRouter are offering it for free right now so I could have used that: https://openrouter.ai/qwen/qwen3.6-plus:free
"[...] In the coming days, we will also open-source smaller-scale variants, reaffirming our commitment to accessibility and community-driven innovation. [...]"
In other words, like GP said, this Qwen3.6-Plus model is not open-weight unlike the other Qwen models.
Almost all means there have been ones before that were not open. So, no contradiction there.
Please send the download link for qwen 3.5-plus.
Also, who cares? If you have the hardware to run a ~400b model i don’t think you count as a home user anymore.
However, my hope is that there will be at least somewhat competitive big and open models as well, from an ethical/ideological perspective. These things were trained on data that was provided by people without their consent, so they should at least be be publicly accessible or even public domain.
I can remember how good Opus 4.5 was. If I'm considering using this, it's most informative to me to compare to the model it's closest to that I have familiarity with.
I'm obviously not switching to this if I want the best model. I'm switching if I'm hopeful that the smaller versions are close to it, or if I want to have more options for providers, or for any other reasons unrelated to getting the highest quality responses possible.
But there are open models that also score 23/25 including Qwen 3.5 27B.
First, try signing up for Z.ai's coding plan. I know how to but I bet you won't be able to.
The absolute disaster that is Z.ai's internet presence shows that these small labs have no ability to market themselves and drive direct sales.
For marketing, they lack capabilities, and releasing open models is the only way for them to remain in the conversation.
For sales, they rely on distribution via OpenRouter, OpenCode etc. Interest with their users is driven by open model performance.
Open sourcing for Chinese labs is not some large national scheme. It is their only way to commercialization.
Only partially tongue in cheek - if it’s not good at marketing itself, that seems like a red flag for capabilities?
What's the issue with signup up for Z.ais coding plan?
- Qwen3.5-Plus
- Qwen3-Max
- Qwen2.5-Max
etc. Nothing really changed so far.
In any case, aside Claude fanboyism, having other plays inch closer to similar performance is always useful. Even if they are "6 months behind" as the pace slows down, this guarantees that there's no huge moat and they'll eventually either get to where the SOTA is, or the difference wont be that big.
I'd rather put fewer eggs in 2-3 big player baskets.
EDIT: Ah, I see. Some kind of promotion. Pretty cool.
It's very generous
I know I could certainly pay anthropic $1500 a day for my use and they'd be delighted... I'd rather pay $0 ... Just a personal preference
As always, we'll have to try and see how it performs in the real world but the open weight models of Qwen were pretty decent for some tasks so still excited to see what this brings.
Ugh, that's not good.
I evaluated Kimi K2 a while back for some text understanding -> summarisation tasks, and of the 100 tasks it hallucinated about 30% of the output. :( :( :(
I guess that it was Kimi K2-Instruct, the first model (or it's fine-tune) in the lineup of Kimi-K2 models. And I remember trying it just for the sake of curiosity, and... except for the almost total absence of the sycophancy and "sugar syrup" in it's outputs, it was not very good at the time. Right now though, if you're still interested in this model family, you could look at Kimi-K2.5 which is way better.
That said, it's still not perfect, and to be honest, looking where things are going with LLMs right now I prefer the use of my own brain (local private inference with power consumption of ~20-25W, having a capability for continuous learning and performing real-world tasks) to the use of any "AI" model (including proprietary models such as Claude 4.6 Opus, Gemini 3.1 Pro and others).
: )
I tried Qwen3.6-plus and it was not as good.
This one seems weird
This means a 100k token request counts the same as a 100-token one. I’ve made about 8000 requests in the last two weeks, averaging around 80k tokens per request. It feels like they’re subsidizing this just to gather data on agentic workflows.
On the downside, the speed is mediocre (15–30 tg/s for GLM-5), and I’ve seen the model glitch or produce broken output about 10 times out of those 8k requests.
you get a generous token limit.
Like Qwen local for it’s privacy, but I trust the privacy of Google/OpenAI/Anthropic more than alibaba.
None should be trusted, unless you are running them locally.
us actually has laws around this and they arent sharing very much with thr us gov today. china shares 100% as required by law. and neither care much about "how long do i cook eggs for", but they do care about code generation a lot.
It's not that, it's about relative risk to your own life. Asking questions about "DEI" for example is much more likely to have adverse effects on your life if you ask Grok or an OpenAI chatbot, though still not that likely.
And the US government has repeatedly shown that it is very interested in collecting all the data available, just like China. In China this is simply done in the open while the US has a veneer of protection for citizens. But where the data collection is forbidden by law they either ignore the law or ask another five eyes member to do the spying and share the results. Both are well documented
I don't know whether you really believe this or it was an off the cuff remark. China is not going to tell you why they plan to arrest you. China is not a benevolent dictatorship.
It doesn't really affect the strength of that particular argument.
And you're being misleading, seemingly on purpose. Please don't.
I feel Alibaba and DeepSeek see themselves more as infra. No urge to control the stack and litigate competition out of existence.