Something is afoot in the land of Qwen
simonwillison.net
simonwillison.net
I've been testing Qwen3.5-35B-A3B over the past couple of days and it's a very impressive model. It's the most capable agentic coding model I've tested at that size by far. I've had it writing Rust and Elixir via the Pi harness and found that it's very capable of handling well defined tasks with minimal steering from me. I tell it to write tests and it writes sane ones ensuring they pass without cheating. It handles the loop of responding to test and compiler errors while pushing towards its goal very well.
This terminology is still very much undefined though, so my version may not be the winning definition.
It's really easy to setup with any OpenAI compatible API and I self host Qwen Coder 3 Next on my personal MBP using LM Studio and just dial in from my work laptop with Zed and tailscale so i can connect from wherever i might be. It's able to do all sorts of things like run linting checks and tests and look for issues and refactor code and create files and things like this. I'm definitely still learning, but it's a pretty exciting jump from just talking to a chat bot and copying and pasting things manually.
I'm aligning on Agent for the combination of harness + model + context history (so after you fork an agent you now have two distinct agents)
And orchestrator means the system to run multiple agents together.
[0] https://www.reddit.com/r/LocalLLaMA/comments/1rivckt/visuali...
It's also driving itself crazy with deadpool & deadpool-r2d2 that it chose during planning phase.
That said, it does seem to be doing a very good job in general, the code it has created is mostly sane other than this fuss over the database layer, which I suspect I'll have to intervene on. It's certainly doing a better job than other models I'm able to self-host so far.
I think this is part of the model’s success. It’s cheap enough that we’re all willing to let it run for extremely long times. It takes advantage of that by being tenacious. In my experience it will just keep trying things relentlessly until eventually something works.
The downside is that it’s more likely to arrive at a solution that solves the problem I asked but does it in a terribly hacky way. It reminds me of some of the junior devs I’ve worked with who trial and error their way into tests passing.
I frequently have to reset it and start it over with extra guidance. It’s not going to be touching any of my serious projects for these reasons but it’s fun to play with on the side.
I can live with this on my own hardware. Where Opus4.6 has developed this tendency to where it will happily chew through the entire 5-hour allowance on the first instruction going in endless circles. I’ve stopped using it for anything except the extreme planning now.
Qwen3.5-35B-A3B means that the model itself consists of 35 billion floating point numbers - very roughly 35GB of data - which are all loaded into memory at once.
But... on any given pass through the model weights only 3 billion of those parameters are "active" aka have matrix arithmetic applied against them.
This speeds up inference considerably because the computer has to do less operations for each token that is processed. It still needs the full amount of memory though as the 3B active it uses are likely different on every iteration.
> Do you feel you could replace the frontier models with it for everyday coding? Would/will you?
Probably not yet, but it's really good at composing shell commands. For scripting or one-liner generation, the A3B is really good. The web development skills are markedly better than Qwen's prior models in this parameter range, too.
What quant are you using? How much ram does it have?
I’m trying to use local models whenever possible. Still need to lean on the frontier models sometimes.
The main quirk I've found is that it has a tendency to decide halfway through following my detailed instructions that it would be "simpler" to just... not do what I asked, and I find it has stripped all the preliminary support infrastructure for the new feature out of the code.
That would seem logical, as the results are then completely deterministic, but it turns out that a suboptimal token may result in a better answer in the long run. Also, allowing for a little bit of noise gives the model room to talk itself out of a suboptimal path.
I wonder if determinism will be less harmful to diffusion models because they perform multiple iterations over the response rather than having only a single shot at each position that lacks lookahead. I'm looking forward to finding out and have been playing with a diffusion model locally for a few days.
For creative things or exploratory reasoning, a temperature of 0.8 lends us to all sorts of excursions down the rabbit hole. However, when coding and needing something precise, a temperature of 0.2 is what I use. If I don’t like the output, I’ll rephrase or add context.
> Blah blah blah (second guesses its own reasoning half a dozen times then goes). Actually, it would be a simpler to just ...
Specifically on Antigravity, I've noticed it doing that trying to "save time" to stay within some artificial deadline.
It might have something to do with the system messages and the reinforcement/realignment messages that are interwoven into the context (but never displayed to end-users) to keep the agents on task.
But how do you see the current thinking level and how do you change it? I’ve been clicking around and searching and adding “effortLevel”:”high” to .claude/settings.json but no idea if this actually has any effect etc.
$ echo 'export ANTHROPIC_EFFORT="high"' >> ~/.zshrc source ~/.zshrc
$ echo 'export ANTHROPIC_EFFORT="high"' >> ~/.bashrc source ~/.bashrc
I prefer settings.json (VSCode) - "claudeCode.environmentVariables": [
{ "name": "ANTHROPIC_MODEL", "value": "claude-opus-4-6" },
{ "name": "CLAUDE_CODE_EFFORT_LEVEL", "value": "high" }
], ...Doh.
If you ask it to do something laborious like review a bunch of websites for specific content it will constantly give up, providing you information on how you can continue the process yourself to save time. Its maddening.
I tend to work on things where there is a massive amount of code to write but once the architecture is laid down, it's just mechanical work, so this behavior is particularly frustrating.
Don't get me wrong, it's very good autocomplete and if you run it in a loop with good tooling around it, you can get interesting, even useful results. But by its nature it is still autocomplete and it always just predicts text. Specifically, text which is usually about humans and/or by humans.
Just think about how many thousands of times you've heard "good morning" after noon both with and without the subsequent "or I guess I should say good afternoon" auto-correct.
So it's not surprising that eventually autocomplete can reach up from those circuits and take on some tasks that have already been made simple enough.
I think what's so interesting is how uneven that reach is. Some tasks it is better than at least 90% of devs and maybe even superhuman (which, in this case, I mean better than any single human. I've never seen an LLM do something that a small team couldn't do better if given a reasonable amount of time). Other cases actual old school autocomplete might do a better job, the extra capabilities added up to negative value and its presence was a distraction.
Sometimes there is an obvious reason why (solving a problem with lots of example solution online vs working with poorly documented proprietary technologies), but other times there isn't. They certainly have raised the floor somewhat, but the peaks and valleys remain enormous which is interesting.
To me that implies there is both lots of untapped potential and challenges the LLM developers have not even begun to face.
Used to have the same thing happening when using Sonnet or Opus via Windsurf.
After switching to Claude Code directly though (and using "/plan" mode), this isn't a thing any more.
So, I reckon the problem is in some of these UI/things, and probably isn't in the models they're sending the data to. Windsurf for example, which we no longer use due to the inferior results.
It's amazing how much foundational prompting and harness matters.
It just decided halfway that, nah, removing the field altogether means you don't have to fix the fallout from making that thing nullable.
Lmao.
That sounds too close to what I feel on some days xD
That's likely coming from the 3:1 ratio of linear to quadratic attention usage. The latest DeepSeek also suffers from it which the original R1 never exhibited.
This is my experience with the Qwen3-Next and Qwen3.5 models, too.
I can prompt with strict instructions saying "** DO NOT..." and it follows them for a few iterations. Then it has a realization that it would be simpler to just do the thing I told it not to do, which leads it to the dead end I was trying to avoid.
I wasn't aware of that, which page mentions that?
I think companies that don’t navigate these correctly eventually lose.
If models like Qwen can get good enough for coding tasks locally, the real shift might be economic rather than purely capability.
Europe really just needs to rally behind Mistral. That's where they should dump their cash.
At the moment my impression is instead that the issue is computational resources. It's important to stay near the frontier though, and to build up ones capacity to train large models.
Consequently I don't think we need Anthropic. It wouldn't be terrible if they came. Especially if they picked a nice location. Barcelona would be very nice, for example.
In any case, there's no way Anthropic's investors in Silicon Valley would countenance such a move.
Also, I'm biased the logical place is Canada, not Europe. Much of the fundamental/foundational research on LLMs, and a large part of the talent, came from universities in Canada anyways.
I think Alibaba needs to just give these guys a blank check. Let them fill it in themselves. Absent that, I'm pretty sure they'll make their own startup.
I do think it'd be a big loss for the rest of the world though if they close whatever model their startup comes up with.
That's very likely to happen once the gap with OpenAI/Anthropic has been closed and they managed to pop the bubble.
I will say we are winning in accessibility. China doesn’t have much of a ramp game
I wonder if you max out your options in China. It seems the Party is suspicious of ambition and high profile winners. I'm sure you can live comfortably, but there's a ceiling.
Isn't it just straight-up illegal in China to refuse the government from using your model? USA isn't perfect, but at least it has active discourse.
Do you imagine an invasion of Taiwan won't involve dropping bombs?
I feel like we should be able to agree that providing authoritarian regimes with high tech tools is immoral in the general case.
If you'd asked me two years ago my answer might have been different.
And to the original point, yeah, I would feel entirely justified in the critique of engineers in providing tools to the US defense apparatus at this point.
At least the Chinese shops are giving their weights away for free, and not demanding that any government ban the rest.
You might even get lucky and someone else does the same. If you manage to learn from their example you might be more competitive in a future round.
I think we are now in the era of oligarchies, and oligarchies maintain power by being highwaymen and extracting tolls, in a kind of rentier capitalist structure.
By throwing LLM models out into the commons, China is disrupting the possibility of this taking hold there.
God bless 'em
Right, because in China if you were criticized for helping the government, the people who criticized you would be in for (probably life-damaging) trouble.
People in Hong Kong died. Over 10,000 were arrested and many are still in prison. The rest are permanently disgraced in their social-credit society.
Again, USA is not perfect, but let's not dream up some fantasy about the CCP.
https://reclaimthenet.org/china-man-chair-interrogation-soci...
But this:
> According to the social credit system, Chinese citizens are punishable if they indulge in buying too many video games, buying too much junk food, having a friend online who has a low credit score, visiting unauthorized websites, posting “fake news” online, and more.
...is just pure bullshit. There were _ideas_ about including these kinds of stuff into the score, but they have never been implemented. At this point, the social credit score is only used to find people who dodge court decisions.
Please ignore the gun pointed at your head / social credit score / masked goons roving about Minnesota / flock cameras / etc as it hasn't been used against you at this point.
If you wish to dispute the veracity of one or more comments in the thread, by all means do so. But please make a substantive argument and (given the nature of the topic) cite sources.
Does that matter? In China people don't judge the state of their civilization by how easily you can insult the police but whether you need to be afraid to meet them on the street. "I can insult my pedophile president" (who doesn't care if you do) isn't exactly a flex.
It does tell us something though that the evaluation of American life now consists of parasocial interactions with the president on social media. I'm starting to belief Bruno Maçães, ex Portuguese secretary of state, was prescient with his diagnosis that American material society has rotted to the point where life is now entirely defined by virtual interactions. That's the difference between China and the US today.
The president's a pedophile, a criminal, undeterred by democracy, economy or social disorder but you can freely yell into the void. Have you considered that in the US one can freely say all these things precisely because that's irrelevant?
Americans will vote for their Congress representatives in November. They will have a chance to decide how they want their government to be run. The US President was already shot-down once by the Supreme Court (tariffs). The system is working. Let the voters decide, and then let it work.
That depends on what's on the ballot.
> The system is working.
If it is, how did you end up here again?
Do you have a legit source for this? When I search for information, I only found this case, “Luo Changqing, a 70-year-old Hong Kong cleaner, died from head injuries sustained after he was hit by a brick thrown by a Hong Kong protester during a violent confrontation between two groups in Sheung Shui, Hong Kong on 13 November 2019.”
None of the other legit sources claim the police killed any of the rioters.
China is bullying lots of countries in the SCS (ramming Philippine coast guard ships, building military installations in the SCS, ...). Not peaceful or responsible.
Many countries in the SCS are doing this. In fact China was late to the game, as Vietnam did it much earlier.
The real political power we have through our vote is probably smaller than the political power most of us here have from the option to quit.
I'm sure it's a very nice place to live if you're content to just stay quiet in society and never put a political sign in your yard or even just talk about the wrong thing with your friend in a WeChat.
Not as bad as China sure, but not as good as other civilized nations.
If you want to put this to the test try crossing the Canadian border and when they ask you the purpose of your visit respond that it's to attend a protest.
Yunseo Chung was not a visitor. She came to the United States from South Korea at age 7. She was arrested last year for peacefully protesting. Charges against her were dropped but the govt. canceled her green card.
The govt. has been trying to deport her since then, but the courts keep blocking it.
https://humanrightsfirst.org/yunseo-chung-v-trump-administra...
While the legality of these actions are being debated in courts, I think most of us can agree that this is reprehensible behavior on part of the Trump admin.
I never claimed to condone the actions of the current admin. The examples of people being deported for protesting that I am familiar with are student visa holders. While I don't personally support the examples that I am aware of, I also recognize that in those specific cases the executive branch appears to be within the bounds of the law. I don't even object to the executive branch having the power to cancel the visas of political dissidents in the general case, merely to how they are choosing to apply it.
It's surprising to me to learn that a green card could be revoked for protected speech. That ought to fall well outside the bounds of the law IMO. Green cards and visas are entirely different things.
It's my understanding that the 1st amendment applies to everyone, not just citizens. So if that's true (not 100% sure about that), how can political speech (protesting) be a valid reason to remove someone from the US?
You can certainly be denied entry for entirely arbitrary reasons. Can you also (as a visa holder) be evicted without notice for same? I think that's generally a safe assumption for any country in the world but would be interested in learning about counterexamples.
Does someone on a short term visa have the protected right to purchase firearms? Visitors aren't even permitted to get a job without the appropriate type of visa. Being allowed to work is a pretty fundamental right.
I expect there's a difference between the bill of rights and the constitution, and likely further nuance as well.
Link? I’m guessing we’re going to see that this definition of “protesting” involves being aggressive and directly in the face of law enforcement officers, not merely holding a sign at a distance.
Please read up on this one example of a US permanent resident. And then justify the actions of the govt against Yunseo Chung.
https://humanrightsfirst.org/yunseo-chung-v-trump-administra...
It just looks a bit ridiculous when students walk out in protest against things that are far outside the influence of their school, city, or even state.
In practical terms, if you're not kind of person who would want to run for an office in the US, China is incredibly comfortable. Cities are safe, with barely any violent crime. Public drug use is nonexistent. And with the US-level AI researcher income, you'd be in the top 0.1% earners.
https://news.ycombinator.com/item?id=47252833
My comment and the linked video says otherwise. The guy was in a private group chat and said some nasty things about the police for confiscating his motorcycle. Now he's arrested and in the Tiger Chair.
How are we explaining this?
Practically, how many care about that? Consider that in other part of the world they also cancel folks based on social media opinion...
and that Benjamin Franklin's opinion on security and freedom? Thats terminally online phenomenon only. I once tried to bring that without specifically mentioned that it came from ol Ben himself to folks IRL. Many thought it was some anarchist blabbers.
In the US people try to hide it and are far more sinister about it, since there are a lot of laws against obvious racism. The cops are also happy in the US to just kill you.
The racism in the US comes out of hate where as what I experienced abroad was more, we don't think you'll fit in and follow the rules and you have to constantly prove that you can.
I didn't spend too much time in China so maybe it is a racist hell hole.
But my experience in Japan was that white immigrants were way more inclined to make a huge deal about the lighter racism they experienced because they had never been somewhere where their skin color was a disadvantage.
I speculate that if you were a permanent minority instead of a visiting inconvenience, then that 'nice' racism you describe would metastasize into the type of racism you see in the USA. It's more friction from time and exposure added on. And, you know, slavery.
Well duh, as recently demonstrated, an US model used by the US gov will 100% end up murdering actual children sooner than later, in this case less than a calendar year in some far flung war that many Americans do not support. Alternatively PRC model used by CCP might kill in some hypothetical future but for national reunification/rejuvenation that many Chinese support. At the end of the day, researchers and population on one side sleeps more soundly.
*edit: not that it matters, but since MAGA can't help but assume, these are all US citizens and green card holders that I am referring to.
To get back to the original point, personally I doubt sentiment on US immigration enforcement would be so significant a deterrent for Chinese talent, who may not share the political views of the American left for whom this is a big concern.
Given the tactics employed by ICE, it's a true shock and horror that most people have more humanity than that.
But I guess a person who can't form a grammatically correct sentence is an example of the sort of people who can rest easy,
https://apnews.com/article/immigration-raid-hyundai-korea-ic...
https://www.koreatimes.co.kr/foreignaffairs/20251112/hundred...
https://www.pbs.org/newshour/nation/attorney-says-detained-k...
The regime is powered by racism and doesn't think through things.
[1] Removals by president: https://www.migrationpolicy.org/article/biden-deportation-re...
No other developed countries have masked goons abducting people in public wearing civil clothes and masks and disregarding every laws of the country (violating private property and foreign embassies, deporting national citizens, and numerous other preposterous bullshit).
Immigration policy enforcement is normal, the madness that has been running in the US for a year isn't.
The concern isn’t IDs exist—it’s who’s demanding them, in what context, and what happens if you can’t comply on the spot.
I forgot that HN is mostly filled with a younger generation that might not get the reference.
The "Papers, please." quote is a common trope in spy movies, books, etc... about the former Soviet Union.
But also, I don't care if it's a tired argument--this isn't about how things are, it's about how we want them to be. I don't want to live in a state action-coerced society.
*Reminder that folks visiting the US on a visa are legally required by the terms of said visa to always carry upon their person at least a copy of identity papers backing up that visa, and that this law has been in place for a very long time.
1.Those who arrived through legal channels (most studied at U.S. universities and remained on H1B visas, with a smaller number through EB5 or other visa categories) and eventualy got green card.
2.Undocumented immigrants, which include several sub-groups/waves. In the 1990s, most came from just a couple provinces, Fujian and southern Zhejiang. After COVID, they were from different parts of China and entered through the southern border.
The contributors to AI development belong to the 1st group. They are spread across the country but a large number work in high-tech companies in Northern California.
The 2nd group was intially concentrated in New York and Southern California (Los Angeles area). Later they have expanded into nearby regions. They provide labor for Chinese-owned small businesses such as restaurants, grocery stores, and hotels.
There is an industry created largely by Chinese political dissidents helping Group 2 through asylum applications using fake materials and exploiting common Western beliefs or narratives about China like human rights concerns. For example, Alysa Liu’s father is an asylum lawyer.
ICE enforcement efforts would likely focus more on Group 2 if they are knowledgable. Ohio should not be a high-priority area. I could be wrong due to changes over time. One indicator you can observe: Are there many Chinese-owned small businesses in your area?
https://www.cato.org/blog/5-ice-detainees-have-violent-convi...
The Bay Area is mostly exempt for now because, after Trump announced ICE was going to surge in SF, a bunch of tech billionaires with economic interests in the region convinced him not to.
Also, over the last year, there have been a bunch of high-profile arrests of Ohioans by ICE. In one example, they arrested someone for showing up to their immigration hearing, leaving their young kid separated from them outside the court.
When I was a deep learning PhD in the first Trump administration, US universities were already very deeply affected by the Muslim ban, and so a lot of talent ended up in other countries.
Sibling commentators are rightfully pointing out that foreigners, especially those who would not be recognized as white, face an onerous and risky customs process with long-term and increasing risks of deportation. When you see a headline like the NIST labs abruptly restricting foreign scientists, _everything_ else feels uncertain. Even if someone doesn't believe they're personally at risk for deportation, they're still seeing everything else.
And then it all boils down to a reputational thing. The era where we were the top choice for research is in the past. If you start a PhD in the US on your resume during this era, you might be anticipating how you'll answe the question of why you weren't good enough to get accepted somewhere better.
Big caveat that I have the perspective of just one US-based former academic.
Besides you can live a comfortable life in PRC nowadays or live in a racist America.
For China, the country, it's a good thing if American AI companies have to scramble to compete with Chinese open models. It might not be massively profitable for the companies producing said models, but that's only a part of the equation
To be honest, it's sort of what I expected governments to be funding right now, but I suppose Chinese companies are a close second.
All is collected in https://imar.ro/~mbuliga/ai-talks.html
Wild times!
If AI could effectively replace people, you wouldn’t need CEOs to keep trying to convince people.
I will also say it’s amusing that the debate is between one and two nines. Neither is objectively great. If you built a system with >3.65 days of downtime in a year that wouldn’t be something you’d brag about in an interview.
Edit: This incident: https://status.claude.com/incidents/kyj825w6vxr8
In any case, two nines of reliability is not impressive.
Probably good to sent alerts early, but they might be going a bit too early.
This is sad for local LLM community. First we lost wizardLM, Yi and others, then we lost Llama and others, now we lost Qwen...
If so, I'm happy that the team held together, and I hope that endogenous tech leads get to control their own career and tech destiny after hard work leads to great products. (It's almost as inspiring as tank man, and the tank commanders who tried to avoid harming him...)
(ducking the downvote for challenging the primacy of equity...)
Is there a better agentic coding harness people are using for these models? Based on my experience I can definitely believe the claims that these models are overfit to Evals and not broadly capable.
They also struggle at translating very broad requirements to a set of steps that I find acceptable. Planning helps a lot.
Regarding the harness, I have no idea how much they differ but I seem to have more luck with https://pi.dev than OpenCode. I think the minimalism of Pi meshes better with the limited capabilities of open models.
This has also been my experience. But isn't the harness sending the instructions on how to invoke a tool? Maybe it is missing the formatting part. What do you think?
But I'll be running this locally for note summarization, code review, and OCR. Very coherent for its size.
I found them to be less than stellar at writing coherent prose. Qwen 3.5 9b was worse in my tests than Gemma 3 4b.
the qwen is dead, long live the qwen.
There would never be an Anthropic/Pentagon situation in China, because in China there isn't actually separation between the military and any given AI company. The party is fully in control.