Running smaller 4B-7B models entirely on the GPU VRAM will get you fast inference, but you will need to scope and define the tasks well. eg, using it the model as a classifier and just feeding it from a queue.
The best performing "agent"-like model to plug into a harness that I have found so far has been Qwen3.6-35B-A3B (mixture of experts) model as I can park most of it in system RAM and CPU, while the VRAM holds the attention/shared weights.
It's definitely workable as a local AI homelab. But expect homelab levels of tuning/fiddling with it.
With the improved support for AMD GPUs I'm finally considering getting a modern 16GB card (and maybe a second one in a few years assuming prices come down)
- https://huggingface.co/unsloth/Qwen3.5-9B-GGUF - https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF (unofficial Qwen 3.8-9B) - https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF (my personnal favorite)
Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.
It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.
The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.
What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?
Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?
The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.
Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think unsloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if well done) the least important information. But with every step you gain less and need to shave of more important things.
> And also be effected by who did the quantization?
My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.
Also, I’d love to use Deepseek directly (or any of the Chinese providers, at that). Seems only fair to pay the lab that built the model. Unfortunately, any requests to Chinese servers is deeply frowned upon here (Belgium, EU). For personal use: sure. As a token intelligence strategy for the company: absolutely fucking not.
Not in my experience. I need to explicitly set my preferred provider(s) for each model, otherwise it bounces me around even within a single session.
You should go to huggingface and maybe create an a/c with a throwaway email and enter your hardware details and that will filter the models for you.
Nothing, really. Might be coming soon, but no.
You probably want to try bonsai, I guess, but don't expect good results.
You could install this yourself for free? I get $0.50 isn’t all that much, but still?
But AI for installing tricky opensource software is indeed a good use case. I do that too.
Even if you spend just 10 minutes of it, I would say $0.50 it's not a bad deal.
One of the dumbest sayings ever. Unless you spend all of your time doing something that makes money, the time is worth $0. You could say that you prefer to do something else during that time and would happily pay to free it up.
I agree, as my time is much more valuable than money.
"You could say that you prefer to do something else during that time and would happily pay to free it up."
Correct, it is applicable to just about everything you pay money for.