No, Qwen 27B isn't the sweet spot
barajas.blog
barajas.blog
1. Mac is slow at prompt processing, the Spark is much better for coding. Most of the time in coding is spent in prompt processing.
2. DFlash addon improves tgen speed 2-3x. (https://huggingface.co/collections/z-lab/dflash)
3. Harness matters, they gave no indication of what they are using. (likely an entire post just to describe a good setup)
The low, double-digit billion param models have improved vastly, notably earlier this year. I'm personally excited for the next iteration later this year because everything keeps getting better. The DFlash addons are a pure computational performance boost, there will be capability improvements later this year in this size range.