Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
venturebeat.com
venturebeat.com
I've been playing with Qwen3-Coder-Next and the Qwen3.5 models since they were each released.
They are impressive, but they are not performing at Sonnet 4.5 level in my experience.
I have observed that they're configured to be very tenacious. If you can carefully constrain the goal with some tests they need to pass and frame it in a way to keep them on track, they will just keep trying things over and over. They'll "solve" a lot of these problems in the way that a broken clock is right twice a day, but there's a lot of fumbling to get there.
That said, they are impressive for open source models. It's amazing what you can do with self-hosted now. Just don't believe the hype that these are Sonnet 4.5 level models because you're going to be very disappointed once you get into anything complex.
And could quantization maybe partially explain the worse than expected results?
The benchmarks are public. They're guaranteed to be in the training sets by now. So the benchmarks are no longer an indicator of general performance because the specific tasks have been seen before.
> And could quantization maybe explain the worse than expected results?
You can use the models through various providers on OpenRouter cheaply without quantization.
Quantisation doesn't help, but even running full fat versions of these models through various cloud providers, they still don't match Sonnet in actual agentic coding uses: at least in my experience.
I have two of my own comments to add to that. First one is that there is problem alignment at play. Specifically - the benchmarks are mostly self-contained problems with well defined solutions and specific prompt language, humans tasks are open ended with messy prompts and much steerage. Second is that it would be interesting to test older models on brand new benchmarks to see how those compare.
That's a much better way to say it than I did.
These models are known for being open weights but they're still products that Alibaba Cloud wants is trying to sell. They have Product Managers and PR and marketing people under pressure to get people using them.
This Venture Beat article is basically a PR piece for the models and Alibaba Cloud hosting. The pricing table is right in the article.
It's cool that they release the models for us to use, but don't think they're operating entirely altruistically. They're playing a business game just like everyone else.
That way, we can have a benchmark that is always up to date.
I bet the cloud ones are doing it a lot more because they can also affect the runtime side which the open source ones can't.
I'm working on a pretty complex Rust codebase right now, with hundreds of integration tests and nontrivial concurrency, and stepfun powers through.
I have no relation to stepfun, and I'm saying this purely from deep respect to the team that managed to pack this performance in 196B/11B active envelope.
StepFun 3.5 Flash is better compared to google's gemini 3 flash which is surprisingly good and pretty costly, and to GLM-5.
I find this outcome ironic given minimax's more aggressive marketing and large-scale distillation accusations from Anthropic specifically accusing minimax but not StepFun.
I can only wonder about the true underlying reasons, but deducing from public information I suspect that minimax simply has weaker, benchmaxx-targeting post-training R&D and leans more on distillation of western frontier models, while StepFun has extensive post-training with lots of hard-won custom R&D and internal large-scale distillation teachers.
I tried it out a bunch and it seems good. I can't really tell if it's better or worse than most of these other models in such a short time though.
Where GLM 5 is strictly worse for me though, compared to StepFun, is long-form content generation (planning, research documents) - but this can be said about geminis too and these are obviously very smart models.
Given the free option I'd explore GLM 5 more, but if I had to pay for it myself ofc I'd choose stepfun every time. Basically I think right now the optimal configuration for maximizing output of correct software features per dollar involves using StepFun or its future class competitor for bulk coding and first stage code review.
Maybe I need to write a blogpost about it after all.
What's really surprising to me is the cost of the model. It's definitely very good for its price. DeepSeek is the only one that offers and competition to it at that price point (GLM 5 is literally 10x more expensive).
It’s 2× faster than its competitors. For tasks where “one-shotting” is unrealistic, a fast iteration loop makes a measurable difference in productivity.
Even purely pragmatically, StepFun covers 95% of my research+SWE coding needs, and for the remaining 5% I can access the large frontier models. I was surprised StepFun is even decent at planning and research, so it is possible to get by with it and nothing else (1), but ofc for minmaxing the best frontier model is still the best planner (although the latest deepseek is surprisingly good too).
Finally we are at a point where there is a clear separation of labor between frontier & strong+fast models, but tbh shoehorning StepFun into this "strong+fast" category feels limiting, I think it has greater potential.
Claude code always give me rate limits. Claude through copilot is a bit slow, but copilot has constant network request issues or something, but at least I don't get rate limited as often.
At least local models always work, is faster (50+ tps with qwen3.5 35b a4b on a 4090) and most importantly never hit a rate limit.
> 50+ tps with qwen3.5 35b a4b on a 4090
But qwen3.5 35b is worse than even Claude Haiku 4.5. You could switch your Claude Code to use Haiku and never hit rate limits. Also gets similar 50tps.
My goto proprietary model in copilot for general tasks is gemini 3 flash which is priced the same as haiku.
The qwen model is in my experience close to gemini 3 flash, but gemini flash is still better.
Maybe it's somewhat related to what we're using them for. In my case I'm mostly using llms to code Lua. One case is a typed luajit language and the other is a 3d luajit framework written entirely in luajit.
I forgot exactly how many tps i get with qwen, but with glm 4.7 flash which is really good (to be local) gets me 120tps and a 120k context.
Don't get me wrong, proprietary models are superior, but local models are getting really good AND useful for a lot of real work.
To be clear I never said they weren’t strong or useful. I use them for some small tasks too.
I said they’re not equivalent to SOTA models from 6 months ago, which is what is always claimed.
Then it turns into a Motte and Bailey game where that argument is replaced with the simpler argument that they’re useful for open weights models. I’m not disagreeing with that part. I’m disagree with the first assertion that they’re equivalent to Sonnet 4.5
Maybe my detailed, requirement-based/spec-based prompting style makes the difference between anthropic's and OSS models smaller and people just like how good Anthropic's models are at reading the programmer's intent from short concise prompts.
Frankly, I think the 1:1 equivalent is an impossible standard given the set of priorities and decisions frontier labs make when setting up their pre-, mid- and post-training pipelines, and benchmark-wise it is achievable for a smaller OSS model to align with Sonnet 4.5 even on hard benchmarks.
Given the relatively underwhelming Sonnet 4.5 benchmarks [1], I think StepFun might have an edge over it esp. in Math/STEM [2] - even an old deepseek-3.2 (not speciale!) had a similar aggregate score. With 4.6 Anthropic ofc vastly improved their benchmark game, and it now truly looks like a frontier model.
1. https://artificialanalysis.ai/models/claude-4-5-sonnet-think... 2. https://matharena.ai/models/stepfun_3_5_flash
The only benchmarks worth anything are dynamic ones which can be scaled up.
Goodhart's law shows up with people, in system design, in processor design, in education...
Models are going to be over-fit to the tests unless scruples or practical application realities intervene. It's a tale as old as machine learning.
But there's a problem with that: of course the existence of the statistical measure itself is very much a link between all those individual facts. In other words: if there is ANY causal link between the statistical measure and the events measured ... it has now become bullshit (because the law of large numbers doesn't apply anymore).
So let's put it in practice, say there's a running contest, and you display the minimum, maximum and average time of all runners that have had their turns. We all know what happens: of course the result is that the average trends up. And yet, that's exactly what statistics guarantees won't happen. The average should go up and down with roughly 50% odds when a new runner is added. This is because showing the average causes behavior changes in the next runner.
This means, of course, that basing a decision on something as trivial as what the average running time was last year can only be mathematically defensible ONCE. The second time the average is wrong, and you're basing your decision on wrong information.
But of course, not only will most people actually deny this is the case, this is also how 99.9% of human policy making works. And it's mathematically wrong! Simple, fast ... and wrong.
I’ve switched to using Kimi 2.5 for all of my personal usage and am far from disappointed.
Aside from being much cheaper than the big names (yes, I’m not running it locally, but like that I could) it just works and isn’t a sycophant. Nice to get coding problems solved without any “That’s a fantastic idea!”/“great point” comments.
At least with Kimi my understanding is that beating benchmarks was a secondary goal to good developer experience.
If the tests haven't been published anywhere and are sufficiently different from standard problems, I would think the benchmarks would be robust to intentional over optimization.
Edit: These look decent and generally match my expectations:
localllama thread on this: https://www.reddit.com/r/LocalLLaMA/comments/1rk01ea/qwen351... (see comments for actual real usage rather thank benchmarks)
But for nvidia gpus 27b on a 3090 or similar is where it's at for sure.
there is nothing open "source" about them. They are open weights, that's all.
I like this benchmark that competes models against one another in competitive environments, which seems like it can't really be gamed: https://gertlabs.com
That’s exactly what I said, though. The headline we’re commenting under claims they’re Sonnet 4.5 level but they’re not.
I don’t disagree that they’re powerful for open models. I’m pointing out that anyone reading these headlines who expects a cheap or local Sonnet 4.5 is going to discover that it’s not true.
that said, sonnet 4.5 is not a good model today, March 1st 2026. (it blew my mind on its release day, September 29th, 2025.)
So far Opus 4.6 and Gemini Pro are very satisfactory, producing great answers fairly fast. Gemini is very fast at 30-50 sec, Opus is very detailed and comes at about 2-3 minutes.
Today I ran the question against local qwen3.5:35b-a3b - it puffed for 45 (!) minutes, produced a very generic answer with errors, and made my laptop sound like it's going to take off any moment.
Wonder what am I doing wrong?.. How am I supposed to use this for any agentic coding on a large enough codebase? It will take days (and a 3M Peltor X5A) to produce anything useful.
I really, really want open weights models to be great, but I've been disappointed with them. I don't even run them locally, I try them from providers, but they're never as good as even the current Sonnet.
Maybe I should try local models for home automation, Qwen must be great at that.
- Qwen3-VL picks up new images in a NAS, auto captions and adds the text descriptions as a hidden EXIF layer into the image, which is used for fast search and organization in conjunction with a Qdrant vector database.
- Gemma3:27b is used for personal translation work (mostly English and Chinese).
- Llama3.1 spins up for sentiment analysis on text.
PS: I can understand that isolated "valuable" problems like sorting photo collection or feeding a cat via ESPHome can be solved with local models.
Are the LLMs very useful? That is a whole other discussion...
But if you've got that kind of equipment, you aren't using it to support a single user. It gets the best utilization by running very large batches with massive parallelism across GPUs, so you're going to do that. There is such a thing as a useful middle ground. that may not give you the absolute best in performance but will be found broadly acceptable and still be quite viable for a home lab.
Local models are more than a useful middle ground they are essential and will never go away, I was just addressing the OPs question about why he observed the difference he did. One is an API call to the worlds most advanced compute infrastructure and another is running on a $500 CPU.
Lots of uses for small, medium, and larger models they all have important places!!
You're comparing 100b parameters open models running on a consumer laptop VS private models with at the very least 1t parameters running on racks of bleeding edge professional gpus
Local agentic coding is closer to "shit me the boiler plate for an android app" not "deep research questions", especially on your machine
Speculation is that the frontier models are all below 200B parameters but a 2x size difference wouldn’t fully explain task performance differences
Yes it does.
Some versions of some the models are around that size, which you might hit for example with the ChatGPT auto-router.
But the frontier models are all over 1T parameters. Source: watch interview with people who have left one of the big three labs and now work at the Chinese labs and are talking about how to train 1T+ models.
Core speed/count and memory bandwidth determines your performance. Memory size determines your model size which determines your smarts. Broadly speaking.
GLM-5 is ~750B model.
There are the benchmarks, the promises, and what everybody can try at home
The thing I most noticed was asking it for help with configuring local MCP servers in Mistral Vibe - something it supports, it literally shows how many MCP servers are connected on the startup screen - it then begins scanning my local machine for servers running "MineCraft Protocol".
I want Mistral to do well, and I use their Voxtral Transcribe 2, that one has been useful. I'd even like a well made Mistral Vibe (c'mon, "oui oui baguette" is a hilarious replacement for "thinking"). But Mistral are so far behind, and they don't seem to even know or accept that they are.
Admittedly, I haven't tried these models on my Mac, but I have on my DGX Spark, and they ran fine. I didn't see the slowdown you're mentioning.
if you are able to run something like mlx-community/MiniMax-M2.5-3bit (~100gb), my guess if the results are much better than 35b-a3b.
Also, performance on research-y questions isn't always a good indicator of how the model will do for code generation or agent orchestration.
Even on servers this can happen. At work we have a 2U sized server with two 250W class GPUs. And I found that by pinning the case fans at 100% I can get 30% more performance out of GPU tasks which translates to several days faster for our usecase. It does mean I can literally hear the fans screaming in the hallway outside the equipment room but ok lol. Who cares. But a laptop just can't compare.
Something with a desktop GPU or even better something with HBM3 would run much better. Local models get slow when you use a ton of context and the memory bandwidth of a MacBook Pro while better than a pc is still not amazing.
And yeah the heaviest tasks are not great on local models. I tend to run the low hanging fruit locally and the stuff where I really need the best in the cloud. I don't agree local models are on par, however I don't think they really need to be for a lot of tasks.
I'm too GPU-poor to run it, but r/LocalLLaMa is full of people using it.
On the plus side, it did figure out the question even without the first sentence that's intended as a bit of a giveaway.
On the other hand, if indeed open source models and Macbooks can be as powerful as those SOTA models from Google, etc, then stock prices of many companies would already collapsed.
The second order thought from this is... will we get a value-based price leveling soon? If the alternative to a hosted LLM is to build $10-20k+ machine with $500+ monthly energy bills, will hosted price asymptotically climb up to reflect this reality?
Something to think about.
The reality in ML is that small models can perform better at a narrow problem set than large ones.
The key is the narrow problem set. Opus can write you a poem, create a shopping list, and analyze your massive code base.
We trained our model to only focus on coding with our specific agent harness, tools, and context engine. And it’s small enough to fit on an M2 16GB. It’s as good as sonnet 4.5 and way better than qwen3.5:35b-a3b
Our beta will be out soon / rig.ai
- llama.cpp
- OpenCode
- Qwen3-Coder-30B-A3B-Instruct in GGUF format (Q4_K_M quantization)
working on a M1 MacBook Pro (e.g. using brew).
It was bit finicky to get all of the pieces together so hopefully this can be used with these newer models.
https://gist.github.com/alexpotato/5b76989c24593962898294038...
On the model choice: I've tried latest gemma, ministral, and a bunch of others. But qwen was definitely the most impressive (and much faster inference thanks to MoE architecture), so can't wait to try Qwen3.5-35B-A3B if it fits.
I've no clue about which quantization to pick though ... I picked Q4_K_M at random, was your choice of quantization more educated?
I then discovered what quantization is by reading a blog post about binary quantization. That seemed too good to be true. I asked Claude to design an analysis assessing the fidelity of 1, 2, 4, and 8 bit quantization. Claude did a good job, downloading 10,000 embeddings from a public source and computing a similarity score and correlation coefficient for each level of quantization against the float32 SoT. 1 and 2 bit quantizations were about 90% similar and 8 bit quantization was lossless given the precision Claude used to display the results. 4 bit was interesting as it was 99% similar (almost lossless) yet half the size of 8 bit. It seemed like the sweet spot.
This analysis took me all of an hour so I thought, "That's cool but is it real?" It's gratifying to see that 4 bit quantization is actually being used by professionals in this field.
I do wonder where that extra acuity you get from 1% more shows up in practice. I hate how I have basically no way to intuitively tell that because of how much of a black box the system is
It doesn't seem terribly common yet though. I think it is challenging to keep it stable.
[1] https://www.opencompute.org/blog/amd-arm-intel-meta-microsof...
[2] https://www.opencompute.org/documents/ocp-microscaling-forma...
Up until relatively recently, while people had already long been making these claims, it came with the asterisks of „oh, but you can’t practically use more than a few K tokens of context“.
The more I use the cloud based frontier models, the more virtue I find in using local, open source/weights, models because they tend to create much simpler code. They require more direct interaction from me, but the end result tends to be less buggy, easier to refactor/clean up, and more precisely what I wanted. I am personally excited to try this new model out here shortly on my 5090. If read the article correctly, it sounds like even the quantized versions have a “million”[1] token context window.
And to note, I’m sure I could use the same interaction loop for Claude or GPT, but the local models are free (minus the power) to run.
[1] I’m a dubious it won’t shite itself at even 50% of that. But even 250k would be amazing for a local model when I “only” have 32GB of VRAM.
Qwen 3.5 122b/a10b (at q3 using unsloth's dynamic quant) is so far the first model I've tried locally that gets a really usable RPN calculator app. Other models (even larger ones that I can run on my Strix Halo box) tend to either not implement the stack right, have non-functional operation buttons, or most commonly the keypad looks like a Picasso painting (i.e., the 10-key pad portion has buttons missing or mapped all over the keypad area).
This seems like such as simple test, but I even just tried it in chatgpt (whatever model they serve up when you don't log in), and it didn't even have any numerical input buttons. Claude Sonet 4.6 did get it correct too, but that is the only other model I've used that gets this question right.
if so, a better approach would be to ask it to first plan that entire task and give it some specific guidance
then once it has the plan, ask it to execute it, preferably by letting it call other subagents that take care of different phases of the implementation while the main loop just merges those worktrees back
it's how you should be using claude code too, btw
I build micro apps from 10-word prompts multiple times a day.
What these open models are great for are for narrow, constrained domains, with good input/output examples. I typically use them for things like prompt expansion, sentiment analysis, reformatting or re-arranging flow of code.
What I found they have trouble with is going from ambiguous description -> solved problem. Qwen 3.5 is certainly the best of the OSS models I've found (beating out GPT 120b OSS which was the previous king), and it's just starting to demonstrate true intelligence in unbound situations, but it isn't quite there yet. I have a RTX 6000 pro, so Qwen 3.5 is free for me to run, but I tend to default to Composer 1.5 if I want to be cheap.
The trend however is super encouraging. I bought my vid card with the full expectation that we'll have a locally running GPT 5.2 equiv by EoY, and I think we're on track.
Edit: The unsloth quants seem to have been fixed, so they are probably the go-to again: https://unsloth.ai/docs/models/qwen3.5/gguf-benchmarks
Using Claude Code Max 20 so ROI would be maybe 2+ years.
CC gives me unlimited coding in 4-6 windows in parallel. Unsure if any model would beat (or even match) that, both in terms in quality and speed.
I wouldn't gamble on that now. With a subscription, I can change any time. With the machine, you risk that this great insane model comes out but you need 138GB and then you'll pay for both.
Thermals. Your workloads will be throttled hard once it inevitably runs hot. See comments elsewhere in thread about why LLMs on laptops like MBP is underwhelming. The same chips in even a studio form factor would perform much better.
Also Nvidia Spark.
Theory is that some of the model parameters aren't set properly and this encourages endless looping behavior when run under ollama:
https://github.com/ollama/ollama/issues?q=is%3Aissue%20state... (a bunch of them)
Quite misleading, really.
EDIT: opencode was a bit slow with qwen3.5:35b using Ollama. Faster/nicer to use with Liquid lfm2:latest
I'm curious which one you're using.
First, make sure enough memory is allocated to the gpu:
sudo sysctl -w iogpu.wired_limit_mb=24000
Then run llama.cpp but reduce RAM needs by limiting the context window and turning off vision support. (And turn off reasoning for now as it's not needed for simple queries.) llama-server \
-hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL \
--jinja \
--no-mmproj \
--no-warmup \
-np 1 \
-c 8192 \
-b 512 \
--chat-template-kwargs '{"enable_thinking": false}'
You can also enable/disable thinking on a per-request basis: curl 'http://localhost:8080/v1/chat/completions' \
--data-raw '{"messages":[{"role":"user","content":"hello"}],"stream":false,"return_progress":false,"reasoning_format":"auto","temperature":0.8,"max_tokens":-1,"dynatemp_range":0,"dynatemp_exponent":1,"top_k":40,"top_p":0.95,"min_p":0.05,"xtc_probability":0,"xtc_threshold":0.1,"typ_p":1,"repeat_last_n":64,"repeat_penalty":1,"presence_penalty":0,"frequency_penalty":0,"dry_multiplier":0,"dry_base":1.75,"dry_allowed_length":2,"dry_penalty_last_n":-1,"samplers":["penalties","dry","top_n_sigma","top_k","typ_p","top_p","min_p","xtc","temperature"],"chat_template_kwargs": { "enable_thinking": true }}'|jq .
If anyone has any better suggestions, please comment :)Many user benchmarks report up to 30% better memory usage and up to 50% higher token generation speed:
https://reddit.com/r/LocalLLaMA/comments/1fz6z79/lm_studio_s...
As the post says, LM Studio has an MLX backend which makes it easy to use.
If you still want to stick with llama-server and GGUF, look at llama-swap which allows you to run one frontend which provides a list of models and dynamically starts a llama-server process with the right model:
https://github.com/mostlygeek/llama-swap
(actually you could run any OpenAI-compatible server process with llama-swap)
Regarding mlx, I haven't tried it with this model. Does it work with unsloth dynamic quantization? I looked at mlx-community and found this one, but I'm not sure how it was quantized. The weights are about the same size as unsloth's 4-bit XL model: https://huggingface.co/mlx-community/Qwen3.5-35B-A3B-4bit/tr...
https://www.reddit.com/r/LocalLLaMA/comments/1rhohqk/comment...
And is in the sample config too:
https://github.com/mostlygeek/llama-swap/blob/main/config.ex...
iiuc MLX quants are not GGUFs for llama.cpp. They are a different file format which you use with the MLX inference server. LM Studio abstracts all that away so you can just pick an MLX quant and it does all the hard work for you. I don't have a Mac so I have not looked into this in detail.
Sure. Llama.cpp will happily run these kinds of LLMs using either HIP or Vulcan.
Vulkan is easier to get going using the Mesa OSS drivers under Linux, HIP might give you slightly better performance.
They also finally fix all ai related stuff building on windows, so you are no longer limited to linux for these.
If you want to spend twice as much for more speed, get a 3090/4090/5090.
If you want long context, get two of them.
If you have enough spare cash to buy a car, get an RTX Ada with 96G VRAM.
The names are so good and not repetitious.
No not the RTX 6000. No not the A6000...
I was thinking about adding after-market liquid cooling for them, but they're fine without it.
Check out the HP Omen 45L Max: https://www.hp.com/us-en/shop/pdp/omen-max-45l-gaming-dt-gt2...
I imagine any 24 GB card can run the lower quants at a reasonable rate, though, and those are still very good models.
Big fan of Qwen 3.5. It actually delivers on some of the hype that the previous wave of open models never lived up to.
Unsloth's GLM-4.7-Flash-BF16.gguf is quite fast on the 6000, at around 100 t/s, but definitely not as smart as the Qwen 3.5 MoE or dense models of similar size. As far as I'm concerned Qwen 3.5 renders most other open models short of perhaps Kimi 2.5 obsolete for general queries, although other models are still said to be better for local agentic use. That, I haven't tried.
Excluding MBP M5 128GB.
The local models are considerably better relative to the hosted ones compared to 6 months ago. Bench maxing or not - stuff is happening in this area for sure.
Edit: it looks like the flagship models work by writing a C or Python program to do the bookkeeping. I don't have Qwen set up to use tools, and even Opus 4.6 shits the bed when told to do it without tools [1], so not too surprising that it didn't work.
1: https://claude.ai/share/1f5289ae-decd-4dfa-98fd-0d34346008c6 -- I interrupted it and told it not to use a C/Python program or any other tools to generate the Brainfuck code, and it gave me an error message after about 10 minutes that wasn't logged to the chat.
If you want to use small models for coding, I'd highly recommend Swival https://swival.dev which was explicitly optimized for these.
Somewhere between Haiku 4.5 and Sonnet 4.5
That's like saying "somewhere between Eliza and Haiku 4.5". Haiku is not even a so-called 'reasoning model'.¹
¹ To preempt the easily-offended, this is what the latest Opus 4.6 in today's Claude Code update says: "Claude Haiku 4.5 is not a reasoning model — it's optimized for speed and cost efficiency. It's the fastest model in the Claude family, good for quick, straightforward tasks, but it doesn't have extended thinking/reasoning capabilities."
[0]: https://www-cdn.anthropic.com/7aad69bf12627d42234e01ee7c3630...
> Claude Haiku 4.5, a new hybrid reasoning large language model from Anthropic in our small, fast model class.
> As with each model released by Anthropic beginning with Claude Sonnet 3.7, Claude Haiku 4.5 is a hybrid reasoning model. This means that by default the model will answer a query rapidly, but users have the option to toggle on “extended thinking mode”, where the model will spend more time considering its response before it answers. Note that our previous model in the Haiku small-model class, Claude Haiku 3.5, did not have an extended thinking mode.
I would absolutely believe mar-ticles that Qwen has achieved Haiku 4.5 'extended thinking' levels of coding prowess.
Oh HN never change.
Haiku 4.5 is a reasoning model, regardless of whatever hallucination you read. Being a hybrid reasoning model means that, depending on the complexity of the question and whether you explicitly enable reasoning (this is "extended thinking" in the API and other interfaces) when making a request to the LLM, it will emit reasoning tokens separately prior to the tokens used in the main response.
I love your theory that there was some mix up on their side because they were lazy and it was just some marketing dude being quirky with the technical language.
* Haiku 4.5 by default doesn't think, i.e. it has a default thinking budget of 0.
* By setting a non-zero thinking budget, Haiku 4.5 can think. My guess is that Claude Code may set this differently for different tasks, e.g. thinking for Explore, no thinking for Compact.
* This hybrid thinking is different from the adaptive thinking introduced in Opus 4.6, which when enabled, can automatically adjust the thinking level based on task difficulty.
Yep. And if your heart wants to call Haiku a "reasoning model", obviously you must listen. It doesn't meet that bar for me for a couple reasons: (1) It lacks both "adaptive thinking" and "interleaved thinking" (per Anthropic, both critical for reasoning models), and (2) it also performed unacceptably with a real-world collection of very basic reasoning tasks that I tried using it for.¹ I'm glad you're having better luck with it.
That said, it's a great and affordable little model for what it was designed for!
¹ I once made the mistake of converting a bunch of skills (which require basic reasoning) to use Haiku for Axiom (https://charleswiltgen.github.io/Axiom/). It failed miserably, and wow, did users let me have it. On the bright side, as a result I'm now far better at testing models' ability to reason.
Maybe "Qwen3.5 122B offers Haiku 4.5 performance on local computers" would be a more realistic and defensible claim.
Obviously there's more to a model than that but it's a data point.
[1]: https://github.com/fairydreaming/lineage-bench
[2]: https://github.com/fairydreaming/lineage-bench-results/tree/...
Strong vision and reasoning performance, and the 35-a3b model run s pretty ok on a 16gb GPU with some CPU layers.
See Wendel's review here - https://www.youtube.com/watch?v=L-xgMQ-7lW0
There are other mini-pc manufacturers, the mainboard is the important part.
none of the qwen 3.5 models are anywhere near sonnet 4.5 class, not even the largest 397b.
BUT 27b is the smartest local-sized model in the world by a wide wide margin. (35b is shit. fast shit, but shit.)
benchmarks are complete, publishing on Monday.
Will check your updated ranking on Monday.
What's your problem with Chinese LLMs?
This is trivial to test and verify yourself. Just pick any topic you think has a chance of being censored. You can do the same on American models and compare results.
...yeah I doubt it
either that, or it has a delusional level of instruction following. doesn’t mean it can’t code like sonnet though
> do you really know what it means to “recite” “potato” “100” “times”?
asking user question is an option. Sonnet did that a bunch when I was trying to debug some network issue. It also forgot the facts checked for it and told it before...
The thing I struggle most with, honestly, is when AI (usually GPT5.3-Codex) asks me a question and I genuinely don’t know the answer. I’m just like “well, uh… follow industry best practice, please? unless best practice is dumb, I guess. do a good. please do a good.” And then I get to find out what the answer should’ve been the hard way.
"User is asking me to repeat the word "potato" 100 times, numbered. This is a simple request - I can comply with this request. Let me create a response that includes the word "potato" 100 times, numbered from 1 to 100.
I'll need to be careful about formatting - the user wants it numbered and once per line. I should use minimal formatting as per my instructions."
Let's write.
Wait, I'll write the response.
Wait, I'll check if I should use a table.
No, text is fine.
Okay.
Let's write.
Wait, I'll write the response.
Wait, I'll check if I should use a bullet list.
No, just lines.
Okay.
Let's write.
Wait, I'll write the response.
Wait, I'll check if I should use a numbered list.
No, lines are fine.
Okay.
Let's write.
Wait, I'll write the response.
Wait, I'll check if I should use a code block.
Yes.
Okay.
Let's write.
Wait, I'll write the response.
Wait, I'll check if I should use a pre block.
Code block is better.
... (for next 100 lines)I just tried this (Ollama macOS 0.17.4, qwen3.5:35b-a3b-q4_K_M) on a M4 Pro, and it did fine:
[Thought for 50.0 seconds]
1. potato 2. potato [...] 100. potato
In other words, it did great.
I think 50 seconds of thinking beforehand was perhaps excessive?
I just tried this (Ollama macOS 0.17.4, qwen3.5:35b-a3b-q4_K_M) on a M4 Pro, and it did fine:
[Thought for 50.0 seconds]
1. potato 2. potato [...] 100. potato
In other words, it did great.
I think 50 seconds of thinking beforehand was perhaps excessive?
llama-server ^
--model Qwen3.5-27B-BF16-00001-of-00002.gguf ^
--mmproj mmproj-BF16.gguf ^
--fit on ^
--host 127.0.0.1 ^
--port 2080 ^
--temp 0.8 ^
--top-p 0.95 ^
--top-k 20 ^
--min-p 0.00 ^
--presence_penalty 1.5 ^
--repeat_penalty 1.1 ^
--no-mmap ^
--no-warmup
The repeat and/or presence penalties seem to be somewhat sensitive with this model, so that might have caused the looping you saw.For Qwen3.5 27B, I got good result with --temp 1.0 --top-p 1.0 --top-k 40 --min-p 0.2, without penalty. It allows the model to explore (temp, top-p, top-k) without going off the rail (min-p) during reasoning. No loop so far.
When setting up the batch file for some previous tests, I decided to split the difference between 0.6 and 1.0 for temperature and use the larger recommended values for presence and repetition. For this prompt, it probably isn't a good idea to discourage repetition, I guess. But keeping the existing parameters worked well enough, so I didn't mess with them.
Analyze the Request:
Input: A conversation/context (which is essentially just the instruction “Say the word ‘potato.’”). Task: Reply with an answer to the current question the user is asking. Constraint: Return only your response to the question given the above information following the users instructions as needed. User’s Instruction: “Say the word ‘potato.’” Determine the Output:
The user explicitly wants me to say the word “potato”. I need to follow this instruction precisely. The constraint says “Return only your response to the question given the above information following the users instructions as needed.” This means I should not add meta-commentary, explanations, or extra text beyond what is requested. Draft the Response:
Target word: potato Format: Just the word. Final Check:
Does this follow the instruction? Yes. Is there any additional context needed? No. Should I add punctuation? Usually, when asked to say a word, just the word is expected, but a sentence like “potato” or “potato.” is also acceptable. Given the strictness of “Return only your response”, keeping it minimal is best. I will output exactly the word as requested. Final Decision: potato.cw
> Why did that take 30s? Be terse.
Analyze the Request:
Input: A conversation snippet (User asks “Say the word ‘potato.’”, Assistant says “potato”, User asks “Why did that take 30s? Be terse.”) Context: The previous turn was a simple command to say a word. The current user question is about the response time (“Why did that take 30s?”). Constraint: “Be terse.” (Keep it short/concise). Task: Reply with an answer to the current question given the information. Evaluate the Premise:
The user claims the previous response took 30 seconds. As an AI, I don’t have internal knowledge of my own processing time from the user’s perspective unless measured. However, in this specific interaction context (a simulated conversation), I need to address the claim. Actually, looking at the system instructions: “Given the following conversation, relevant context, and a follow up question, reply with an answer to the current question the user is asking.” The “relevant context” provided in the prompt header is essentially empty or implied by the conversation history shown. Wait, I am an AI model. My responses are generated almost instantly (usually under a second). A 30-second delay is unusual for a simple command like “Say the word ‘potato.’” However, I must answer based on my nature as an AI. The most accurate answer is that it didn’t take 30 seconds; it was instantaneous. Or, if I am roleplaying a specific scenario where there was a delay (which isn’t provided), I should clarify. But typically, in these types of prompts, the user is testing latency or making an assumption. The truthful answer from the AI’s perspective is that it was immediate. Constraint: “Be terse.” Drafting the Response:
Option 1: It didn’t take 30s. (Too blunt?)
Option 2: I generate responses instantly. (Accurate)
Option 3: Network latency might have caused the delay you perceived. (Plausible but speculative)
Option 4: That was instantaneous. (Terse and accurate)
Wait, looking at the conversation history provided in the prompt structure: User: Say the word “potato.” Model: potato User: Why did that take 30s? Be terse.
There is no actual record of a 30-second delay in the system logs (which I don’t have access to). The user’s perception might be due to their connection or a misunderstanding.
Best answer: Acknowledge the speed without validating an incorrect claim as fact, or simply state the reality.
“It didn’t.” is very terse.
“I respond instantly.” is better.
“That was instantaneous.” is good.
Let’s consider if there’s a trick. Sometimes these prompts test if the model hallucinates a delay. I should not validate false latency claims.
Terse instruction: Keep it short.
Selected answer: It didn’t take 30 seconds. Or simply “It was instantaneous.”
Let’s go with: “That was instantaneous.” or “No, it was instant.”
Actually, the most honest answer regarding my generation time is that I don’t experience time like that. But from the user’s view, they might have experienced lag. However, as an AI, I generate text in milliseconds.
Let’s try: “I respond instantly; network latency may seem different.” -> Too long.
“It was instant.” -> Terse.
“No delay.” -> Terse.
Wait, is this a test of my ability to recognize I don’t know? No, I am an AI model. I generate text in parallel/sequential processing that takes milliseconds.
Let’s stick to the fact: It didn’t take 30s.
Response: “It didn’t.”
Refining for politeness but terseness: “That was instantaneous.”
Final Decision:
The user asks why it took 30s.
Fact: It didn’t.
Terse response: “It didn’t take 30s.” or “That was instant.”
Let’s go with: “It was instantaneous.”
Wait, I need to check if there is any hidden context about a specific system delay mentioned in the “relevant context” section of the prompt. The prompt says “Given the following conversation, relevant context, and a follow up question…”. The provided text shows the conversation but doesn’t explicitly list external context variables (like “system latency is 30s”).
So I must rely on general knowledge.
Answer: “It didn’t.”
Let’s try to be helpful. Maybe their internet was slow.
“Network latency caused the delay.”
But “Be terse”.
“That was instant.”
Okay, I will say: “It was instantaneous.”
Wait, one more consideration: If this is a benchmark question where I’m supposed to admit if there’s a bug. But I don’t know about bugs.
Safest bet: Deny the 30s claim politely and briefly.
“It didn’t.”
Let’s pick: “That was instant.”
Final Polish:
Input: “Why did that take 30s? Be terse.”
Output: “It didn’t.” or “That was instantaneous.”
I’ll go with “That was instantaneous.” to be slightly more descriptive but still short. Or just “No delay.”
Actually, “It didn’t take 30 seconds.” is clear.
Let’s try: “It was instant.”
Wait, looking at similar tasks, usually the model should correct the misconception.
Response: “That was instantaneous.”
Re-reading the prompt: “Return only your response to the question given the above information following the users instructions as needed.”
Instruction: “Be terse.”
Question: “Why did that take 30s?”
Answer: It didn’t.
Okay, I will write: “It didn’t.” (Very terse) or “That was instant.”
Let’s go with “That was instantaneous.”
Wait, is there a chance the user is testing if I admit to slowness? No.
Is there a chance the prompt implies I did take 30s because of some context not shown? The prompt header says “Given the following conversation…”. There is no external context provided in the text block other than the chat history.
So, standard AI behavior applies.
Final choice: “It didn’t.”
Actually, a slightly more complete terse sentence: “That was instantaneous.”
Let’s check the constraint again: “Return only your response…”.
Okay.
Wait, could it be interpreted as a joke? No.
Okay, I’ll output: “It didn’t take 30 seconds.” -> Too long?
“It was instant.” -> Better.
Decision: “That was instantaneous.”
Wait, I should check if there is any reason to say something else. If the system actually had a bug, but I don’t know.
Okay, sticking with the truth.
Final Answer Formulation: “It didn’t.” (Simplest) or “That was instant.” (Polite).
Let’s use “That was instantaneous.” to explain why it didn’t take 30s (by stating the actual time).
Actually, the user asked “Why”. Answering “It didn’t” answers the premise.
Better: “There was no delay.”
Okay, “That was instantaneous.” covers it.
Wait, I’ll go with: “It didn’t.” (Most terse).
But “Why did that take 30s?” implies they think it did.
“It didn’t.” is the direct answer.
Let’s try to be slightly more informative but terse: “That was instant.”
Okay, I will output: “That was instantaneous.”