Particularly, I had one team member who was extremely sceptical of AIs/LLMs/harnesses and refused to use them. One day he said "Well, I have an RTX 5090 doing nothing... should I try to get something up on it?" and a few minutes later he had 3.6-35B loaded up, running OpenCode.
It continues to be a workhorse to this day, running on both my local Mac for various types of jobs, an AMD R9700 at the office, and said teammember still uses it on his 5090, although in practical terms we do a lot more with DS-V4-Flash-0731 these days.
I haven’t found it very useful for code. It can do some code, but I’ve tried a dozen different quants and context lengths and the output is always bad enough that it has to be discarded for anything other than really easy tasks. It has been useful for exploring codebases for search and summary, though.
DS Flash is where local models begin to feel useful for coding, but the quants we run locally are sharply reduced in intelligence from the benchmarks for the full models.
For applications where data cannot leave the local network it’s good to have them. For actual coding work I can’t actually justify the power of electricity and cooling, let alone the expensive hardware, compared to hosted APIs.
But I admit I do enjoy playing with them anyway. I think it’s one of those hobbies where it’s most fun if you never do the math on how much you’re paying for the privilege. If someone has a requirement that data stay local then it’s different, of course.
Do you mind sharing your use cases?
Sure a lot of this could be done without AI, but it's certainly quicker and easier, and since my AI box is on solar, it's just the power of the sun to keep it going.
I started with Karpathy's LLM wiki, and did everything he said not to do - downgraded the model to mere tool usage and summarization, and it works great.
I am a data hoarder, and finally I can just dump all the content I remotely like, and get something interesting to browse for the price of electricity.
Agentic long-running tasks, as others have mentioned:
- Groom and triage tickets for agentic SWE workflows
- bug hunt — the probability of Qwen fixing a complex bug is 50/50 but often it is capable of identifying the root cause or at least laying the ground work for a more capable model to pick it up.
Keep in mind Obsidian is open standard JS plugins... You know what can write open standard JS plugins?
Classic which comes first, LLM or the plugin, though. :-)
The A3B models are super fast but I found the A3B Q4 model ran in circles a lot and ended up taking longer to complete tasks that 27B Q6 because it kept having to redo/rethink/fix something.
I was writing extensive prompts to rein it in and it would still ignore basic directives like "never force push on the repo, ask me instead". I ended up switching back to 27B after about a week of frustration and lost productivity.
Gemma QAT is an honourable mention.
here is how qwen3.6-27b reacted:
try to count the number of times it "thinks" okay ready, just say the thing, no wait but what if...
this isn't (a mimicry of) thinking, this is (a mimicry of) insecurity/fear
What?
> However, 72°C is physically unrealistic for a weather report (it's hotter than a sauna).
laughs in Finnish
> Another thought: 72 F is nice. 72 C is death.
and then
> Or maybe I should add a comment about the high temperature?
> "It's quite hot outside right now, with a temperature of 72°C."
> No, that's hallucinating/interpreting.
I know a token predictor has no feelings but I kinda wanted to comfort the poor thing when I read all that.
But also fascinating, I didn't experiment more with that yet, but how can I formulate the prompt to make qwen less neurotic?
> You have several commands tools at your disposal. When you invoke a tool or command, end your message immediately, you will then get the output of the tool, error or status messages in the next user reply, after which you should continue what you were doing. Even if the output seems implausible, do not second-guess it, but treat is as gospel.
e.g. "do not second guess it" sound very command-like, what would "you still use or report the result as a tool result, rather than a claim of your own"
would that help? Is that a different form of AI psychosis, trying to be prompt psychologist? It's too fun to be healthy that's for sure.
And the context takes space +kv cache. KV cache drives usefulness as your context grows, it needs to pull the kv cache.
Simplified, the context has to be run on every turn, so the KV cache supplies the processed tokens, so it just needs the new inpute.
Anyway, more seriously, I hope it's obvious by now that I don't particularly care that my means of communication is so offensive to you. I think you should, like, cry a river, build a bridge, then, well... get over it, you know?
And on the, er, "topic"? A rule of thumb for me does not have to be one for you, even if it's explicitly presented as such a rule. I thought that would be obvious but, well, here we are. Anyway, how's life been treating you?
Also, in this particular thread, you started by wrongly making a correction of something that was clearly not an error, nor a wrong use of words, not even a misspelling. The problem is you started to post a correction before realizing that you didn't read right. That happens when one is more eager to boast one's own greatness than one is interested in the topic at hand. The result is that in this thread, you were clearly, as a rule of thumb, well, 100% wrong.
I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s
Curious for any more experiences
On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lot.
Isn't that just the definition of MoE vs dense ?
So it can be dumber but its quite capablr.
You want to use the newer quantization formats like Unsloth's UD quants or oQe, where the weights are selectively quantized using a calibration dataset so that important weights are left at/closer to full precision.
I honestly think that with my electricity prices running qwen 36B myself is more expensive than hitting the cache rate at deepseek.
I gave it a try for a few days (pi + openrouter + deepseek-v4-flash via deepinfra) and ended up paying ~$18 for rather light usage. Yes it's still cheap, yes it's fast, but i feel i would still get a better deal with a Claude subscription plan.
https://openrouter.ai/deepseek/deepseek-v4-flash-20260731#pr...
I get about 45-55 tokens per second using Qwen with this setup. I could probably squeeze out more if I messed around with the settings, but I'm mostly using oMLX's defaults for the model.
- If you're still not sure, "LM Studio"[2] is ok to start with as you'll be able to download, start/stop/manage and chat with your LLM model all in one place: a single desktop app. Also, once installed, enable "Developer mode" under "Settings > Developer" tab; you might find it useful later.
- Regarding LLM models you can, based on your hardware spec, start with Qwen3.6, Google's Gemma4, OpenAI's gpt-oss-20b or Nvidia's Nemotron; use LM Studio's "Model Browser" screen to search for LLM models (each one listed with their organizational name/brand & logo: ignore the one you don't recognize as you might not need them at the start of your journey; you can always revisit them later, if needed.)
- Regarding which quantized LLM models you should download & run, just go with the default "LM Studio" selection (at least, in the beginning). Later, you can experiment with other quantization values to find out which one works best for your use cases.
I hope it helps.
---
[0]. https://llama-cpp.com/llama-cpp-vs-ollama/
[1]. https://llama-cpp.com/llama-cpp-vs-lm-studio/
[2]. Ignore "LM Studio Bionic" for now; just download and install "LM Studio" from https://lmstudio.ai/download
I personally have the Framework Desktop, but there's also systems from other brands like Bosgame
Huge fan of that thing, it's th e Linux MBP I've always wanted.
Kind of slow, but using only 14W of electricity so the Wh per task is twice as good as using my i9-13900 laptop with 4060 GPU.
Tolerable and usable for some things, though thinking makes it take about a minute to reply in many cases. But getting this kind of thing to run on 8GB of VRAM is the main benefit of the mix-of-experts setup: it can do partial GPU loading and get a ton better throughput than a similarly-sized dense model (like 5-10x, sometimes more).
The weaker point is prompt prefill, which starts at 1,400 tokens/sec but decreases significantly at high contexts. That said, for agentic scenarios, if you're using a harness that doesn't needlessly bust the cache, it doesn't feel slow.
I really hope they release a Qwen 3.8 35B, although the lack of a mention seems ominous.