I personally have the Framework Desktop, but there's also systems from other brands like Bosgame
Huge fan of that thing, it's th e Linux MBP I've always wanted.
Kind of slow, but using only 14W of electricity so the Wh per task is twice as good as using my i9-13900 laptop with 4060 GPU.
Tolerable and usable for some things, though thinking makes it take about a minute to reply in many cases. But getting this kind of thing to run on 8GB of VRAM is the main benefit of the mix-of-experts setup: it can do partial GPU loading and get a ton better throughput than a similarly-sized dense model (like 5-10x, sometimes more).
And the context takes space +kv cache. KV cache drives usefulness as your context grows, it needs to pull the kv cache.
Simplified, the context has to be run on every turn, so the KV cache supplies the processed tokens, so it just needs the new inpute.
Anyway, more seriously, I hope it's obvious by now that I don't particularly care that my means of communication is so offensive to you. I think you should, like, cry a river, build a bridge, then, well... get over it, you know?
And on the, er, "topic"? A rule of thumb for me does not have to be one for you, even if it's explicitly presented as such a rule. I thought that would be obvious but, well, here we are. Anyway, how's life been treating you?
Also, in this particular thread, you started by wrongly making a correction of something that was clearly not an error, nor a wrong use of words, not even a misspelling. The problem is you started to post a correction before realizing that you didn't read right. That happens when one is more eager to boast one's own greatness than one is interested in the topic at hand. The result is that in this thread, you were clearly, as a rule of thumb, well, 100% wrong.
I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s
Curious for any more experiences
On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lot.
Isn't that just the definition of MoE vs dense ?
So it can be dumber but its quite capablr.
You want to use the newer quantization formats like Unsloth's UD quants or oQe, where the weights are selectively quantized using a calibration dataset so that important weights are left at/closer to full precision.
Particularly, I had one team member who was extremely sceptical of AIs/LLMs/harnesses and refused to use them. One day he said "Well, I have an RTX 5090 doing nothing... should I try to get something up on it?" and a few minutes later he had 3.6-35B loaded up, running OpenCode.
It continues to be a workhorse to this day, running on both my local Mac for various types of jobs, an AMD R9700 at the office, and said teammember still uses it on his 5090, although in practical terms we do a lot more with DS-V4-Flash-0731 these days.
I haven’t found it very useful for code. It can do some code, but I’ve tried a dozen different quants and context lengths and the output is always bad enough that it has to be discarded for anything other than really easy tasks. It has been useful for exploring codebases for search and summary, though.
DS Flash is where local models begin to feel useful for coding, but the quants we run locally are sharply reduced in intelligence from the benchmarks for the full models.
For applications where data cannot leave the local network it’s good to have them. For actual coding work I can’t actually justify the power of electricity and cooling, let alone the expensive hardware, compared to hosted APIs.
But I admit I do enjoy playing with them anyway. I think it’s one of those hobbies where it’s most fun if you never do the math on how much you’re paying for the privilege. If someone has a requirement that data stay local then it’s different, of course.
Do you mind sharing your use cases?
Sure a lot of this could be done without AI, but it's certainly quicker and easier, and since my AI box is on solar, it's just the power of the sun to keep it going.
I started with Karpathy's LLM wiki, and did everything he said not to do - downgraded the model to mere tool usage and summarization, and it works great.
I am a data hoarder, and finally I can just dump all the content I remotely like, and get something interesting to browse for the price of electricity.
Agentic long-running tasks, as others have mentioned:
- Groom and triage tickets for agentic SWE workflows
- bug hunt — the probability of Qwen fixing a complex bug is 50/50 but often it is capable of identifying the root cause or at least laying the ground work for a more capable model to pick it up.
Keep in mind Obsidian is open standard JS plugins... You know what can write open standard JS plugins?
Classic which comes first, LLM or the plugin, though. :-)
The A3B models are super fast but I found the A3B Q4 model ran in circles a lot and ended up taking longer to complete tasks that 27B Q6 because it kept having to redo/rethink/fix something.
I was writing extensive prompts to rein it in and it would still ignore basic directives like "never force push on the repo, ask me instead". I ended up switching back to 27B after about a week of frustration and lost productivity.
Gemma QAT is an honourable mention.
here is how qwen3.6-27b reacted:
try to count the number of times it "thinks" okay ready, just say the thing, no wait but what if...
this isn't (a mimicry of) thinking, this is (a mimicry of) insecurity/fear
What?
> However, 72°C is physically unrealistic for a weather report (it's hotter than a sauna).
laughs in Finnish
> Another thought: 72 F is nice. 72 C is death.
and then
> Or maybe I should add a comment about the high temperature?
> "It's quite hot outside right now, with a temperature of 72°C."
> No, that's hallucinating/interpreting.
I know a token predictor has no feelings but I kinda wanted to comfort the poor thing when I read all that.
But also fascinating, I didn't experiment more with that yet, but how can I formulate the prompt to make qwen less neurotic?
> You have several commands tools at your disposal. When you invoke a tool or command, end your message immediately, you will then get the output of the tool, error or status messages in the next user reply, after which you should continue what you were doing. Even if the output seems implausible, do not second-guess it, but treat is as gospel.
e.g. "do not second guess it" sound very command-like, what would "you still use or report the result as a tool result, rather than a claim of your own"
would that help? Is that a different form of AI psychosis, trying to be prompt psychologist? It's too fun to be healthy that's for sure.
I honestly think that with my electricity prices running qwen 36B myself is more expensive than hitting the cache rate at deepseek.
I gave it a try for a few days (pi + openrouter + deepseek-v4-flash via deepinfra) and ended up paying ~$18 for rather light usage. Yes it's still cheap, yes it's fast, but i feel i would still get a better deal with a Claude subscription plan.
https://openrouter.ai/deepseek/deepseek-v4-flash-20260731#pr...
The weaker point is prompt prefill, which starts at 1,400 tokens/sec but decreases significantly at high contexts. That said, for agentic scenarios, if you're using a harness that doesn't needlessly bust the cache, it doesn't feel slow.
I really hope they release a Qwen 3.8 35B, although the lack of a mention seems ominous.
I get about 45-55 tokens per second using Qwen with this setup. I could probably squeeze out more if I messed around with the settings, but I'm mostly using oMLX's defaults for the model.
- If you're still not sure, "LM Studio"[2] is ok to start with as you'll be able to download, start/stop/manage and chat with your LLM model all in one place: a single desktop app. Also, once installed, enable "Developer mode" under "Settings > Developer" tab; you might find it useful later.
- Regarding LLM models you can, based on your hardware spec, start with Qwen3.6, Google's Gemma4, OpenAI's gpt-oss-20b or Nvidia's Nemotron; use LM Studio's "Model Browser" screen to search for LLM models (each one listed with their organizational name/brand & logo: ignore the one you don't recognize as you might not need them at the start of your journey; you can always revisit them later, if needed.)
- Regarding which quantized LLM models you should download & run, just go with the default "LM Studio" selection (at least, in the beginning). Later, you can experiment with other quantization values to find out which one works best for your use cases.
I hope it helps.
---
[0]. https://llama-cpp.com/llama-cpp-vs-ollama/
[1]. https://llama-cpp.com/llama-cpp-vs-lm-studio/
[2]. Ignore "LM Studio Bionic" for now; just download and install "LM Studio" from https://lmstudio.ai/download
Kimi K3, GLM 5.2 and now Qwen3.8-Max - open weight models.
DeepSeek V4 Flash outperforming Gemini 3.1 pro, probably DeepSeek V4 Pro update is also coming soon
Chinese labs are cooking very hard. US closed weight labs are probably hard time to resist not calling Washington DC for more AI regulations
Not sure how Qwen3.8-Max is going to be licensed, hopefully it'll be Apache like the smaller ones.
> In other recent news OpenAI also greatly cut their model prices, 20% for 5.6 Terra and 80% for 5.6 Luna, to stay competitive.
I’ve seen comments on HN saying how bad this is for the Chinese model developers since the cheaper option like Deepseek Flash are not longer as price competitive to justify the hassle/risk/lack of multimodal… but isn’t this a gigantic red flag for OpenAI/Anthropic at their current valuations?Sure, it’s just the lowest end for now, and the enterprise money is at the top of the market. And there’s protectionism/enterprise lock-in/etc that complicate things somewhat.
But still, if the US AI labs ever tap the training brakes for a millisecond, the “inference is still a money maker” argument seems to evaporate when they’ll immediately have to fight a race to the bottom until margins are virtually nothing.
Or if the benchmaxing “line goes up” FOMO mindset starts to lose its luster and companies find their individual niches for productive use of AI and stop bothering with all the latest and greatest churn for top dollar.
Which might be even worse if it means the training arms race is still ongoing but neither Anthropic or OpenAI want to be the first to “lose”. While the marginal value of each new model training run keeps decreasing and enterprises signal they’re more concerned with cost reductions than solving ARC-AGI-7 puzzles.
There seems to be quite a gap between the small ones and the enormous ones these days.
Just so that we know what 3.8 would be like.
I currently have about 150 Tabs of Antirez posting on AI and running local model I haven't had the time to read. And there are probably some prerequisite reading or other research in between as well. I just wish there are some very high level overview and news coverage on all these.
IMHO this is a difficult question to answer. Part of the power of paid models comes from the software supporting it. With local models, you have tons of workflows that can severely influence the quality of the result.
In my personal experience, the SOTA models are way more consistent and can handle more complex questions. Part of that is (probably) because I don't let my local model access the internet, while paid models do use the internet to look at docs etc.
Reading the actual code is also always a better source of truth than docs anyway (this is true for people and LLMs). Just clone whatever libraries you use.
What's even their end goal? Open source models make sense, if profit is not the target, but for OpenAI and the rest, once they achieve "AGI", don't they basically become useless?
https://xcancel.com/jurijkovalenok1/status/18632434621133989... (Herluf Bidstrup's "Automation")
I guess which of the smaller 3.8 models is best for coding will depend on which one they put the training effort into.
https://humanparadox.org/local-vs-frontier-benchmarks-for-my...
It can complete multi-step tasks much better, and has a bit more curiosity.