Thanks for flagging, this is on my local rig and it's driving my display too. I'm curious now, will take a closer look. These are the tok/s as reported by LMStudio.
EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b
EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b
EDIT: I get 259 tok/s with the Q4_K_M quant
https://dl.jszym.com/share/boards/pictures/Screenshot_202607...