With this model size I've found that the harness seems to matter more. I've moved on to little-coder rather than raw pi with qwen3.6 27b personally, it might be worth taking a look.
Model Adapter Suite Score Passed Tasks
--------------------------------- ------------ -------------- ------ ------ -----
local/ornith-1.0-35b little_coder aider_polyglot 36.0% 81/225 225
local/ornith-1.0-35b pi_devstack aider_polyglot 39.6% 89/225 225
local/ornith-1.0-35b pi_vanilla aider_polyglot 32.0% 72/225 225
Little Code does a little better than raw Pi, although maybe not better than my personal Pi setup: https://github.com/lhl/devstack