821 karma · joined September 19, 2011
Why not buy a Low End Box and run one of the WG Setup scripts to get you going? You'll lose the anonymity, but its your box, a clean IP, significantly cheaper than a commercial vpn.
Not completely true. It's memory AND memory bandwidth. You can have 1tb of memory but if you have awful memory-bandwidth you'll also have slow tok/s. A3B helps with this, but so does MTP.
From my experience, you'd be better off running the dense 27b-mlx with MTP than the 3.6 version with A3B. You say your model is ~20GB of ram, but the 3.8:27b-mlx is 18GB and gets me very reasonable tok/s, and greater speed if you disable thinking when not required.
Qwen's advances do (currently) have merit.
Surely you don't want them to be the reason the bubble bursts?
Edit: Having used qwen3.8:27b-mlx on MBP M4 64GB, I get around ~45 tok/s. A3B would be great for smaller devices, but it's definitely usable. As I understand it it's a mixture of MLX and MTP.
The main improvements lately were:
- line/function prediction
- Claude Code plugin
And I can use Claude Code in any terminal. So I've switched to zed.dev and use Claude there. Wish it had the line prediction though.
Now days, I don't have Facebook, I don't play games, and the only forum I call home is this one. Times have changed, but so have I. At least I can reminisce on the good times.
on 64GB M4 I find it's able to do things fairly well. The few times I run out of tokens, I hop over to that and I'm mostly unimpeded. I compare it to the Haiku models, where you have to go in and be surgical about your changes, or like others have said, guide a junior.
on 32GB M5, I find that it works, but around the 30% ctx threshold it slows down quite substantially, so more need to be surgical in your requests. I'll often just have my IDE open and Claude. But maybe I've been too comfortable talking to Sonnet/Opus and so forget I need to be more deliberate in my requests.
My finding here is that the harness is a big part of the problem. CC seems to be very good with Qwen in my experience. Better than OpenCode.
I also run DeepSeek for some other non-structured data tasks and to generate a to-do out of that. That's not coding, so won't go into that, other than to say it's very competent as a small model left to run in the background and automate small parts of my life and process.
tl;dr it's totally doable on a 32gb mbp using ollama, but be precise in your requests and guidance.