Have you tried 27B class models like qwen3.6?
Do you have a set of models you'd like me to look at?
I just saw this simple patch to enable MTP (potentially 2x performance) on older GPUs (Kepler etc), so maybe it will work for you
https://github.com/ggml-org/llama.cpp/pull/25680
Also, for Qwen, the 4 bit _XL quantization seems to have a good balance of performance to size.
Also if nothing else the below project lets you use an NVidia graphics card as low-latency swap, which has been nice as a buffer as RAM prices remain high and leaves me eyeing that 24GB card you mentioned as an alternative...