It did take a little while for Unsloth to update the Laguna S 2.1 quants to fix the yarn_attn_factor, and so it was a bit frustrating getting that quantization running right, but almost always, I pick the unsloth quantization if there is one. (Still waiting/hoping for a Ling 3.0 Flash.)
They are the most reliable in my experience, but if you have alternatives you trust I'd love to know
I have had the same experience with gemma 4 on same tasks being refused. But this is when working with cyber offensive tasks and the like. It excels in coding and is very fast on consumer hardware. So I would say use the right tool for the right task.
What uncensored models can you recommend?
Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes works, but it actually screws up reasoning sometimes
No issues with llama.cpp.