I question this choice. Coding models have improved significantly in the past 2 years.
I would try using ~1/2 your available ram and iterate from there.
If you have 32GB of RAM, I would give Qwen 3.8 a try. All you would have to do is run "ollama pull qwen3.8:27b" then "ollama run qwen3.8:27b". If you have 16GB of RAM, I would try Gemma 4.
[1] ollama.com
- Linux, general use case, balance - 16 GB RAM - 10 GB VRAM
Recommendation: Kimi-K3
This checks out.
64 or 128GB RAM, 6 or 64GB of VRAM....
Not sus at all.
Kimi-K3
I think the slop site is broken or compromised...
fitmyllm is much better
By the way own exploration sort of led me to qwen3.5:9b for the best case scenario balanced model considering almost 10-11GB of RAM is almost always gone anyway. Even with aggressive app quitting.
For my anemic 6GB built-on 14GB Qwen seems to be the best bet, not great reviews but from my limited testing its pretty impressive.