Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on/off), it seems to me like Flash either is less "token efficient" or just likes to think more, or I'm doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.
Most of the Chinese models get into long complicated thinking loops before they accomplish something more complicated.
This is also true for Deepseek 4(.1) .
Any lower (or broken) quantization could do that, not what I'm talking about though. They work fine for their size, as far as I can tell. Just surprised the Flash would reason for longer than the Pro.
I prefer GLM 5.3 (Flash or not) over Deepseek 4.1 Flash because the GLM models are significantly less verbose.
curious why the HF pill (on the right) always has inaccurate values
I noticed the same, and I wonder as well.
I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
Packed 4-bit weights are often identified as u8 byte arrays in the safetensors metadata, so the HF UI counts them as half a weight each.
I would think there is sufficient information in the various config files to work this out, based what I've seen in my own quant artifacts.
Yeah, I agree it's probably fixable, but I think a naive interpretation of the metadata is probably the source of this bug (which HF has had for as long as I can remember)