also can you use it for fine tuning?
also can you use it for fine tuning?
A model similar in size to Laguna S 2.1, but with only 6B active parameters, should be a notable amount faster, so I would imagine 25-30 t/s would be a reasonable guess for where Qwen 3.8 Flash Next will land.
DFlash2 might improve all these numbers. It wasn't available last I was testing new models on the Strix Halo; I've only used MTP (which doesn't generally improve MoE models, but I believe DFlash2 can).
Given software improvements, I'm hopeful an MoE in this size range will be the sweet spot that pushes past 40 t/s and is also smart enough for real work. Qwen 3.8 27B is finally a self-hostable model that's smart enough, but it thinks so hard it still isn't really useful for agentic interactive use.
Note also prefill with large models is pretty slow on the Strix Halo (300 t/s, maybe). Time to first token is a painful wait, when using it interactively with large models.
I am using the PrismaAQUA
standard 9.7 t/s
+ Dflash2 30 t/s
+ torch-compile 37 t/s
c8 = 177 t/s
Also, 4-bit has measurable intelligence loss. Sometimes worth it, but, at this size models are barely smart enough at 8 or 6.
The output quality is higher. It's held at full precision (not quantized).
These are roughly the settings I use: https://github.com/kyuz0/amd-strix-halo-toolboxes#kernel-par...