> Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc?
Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think unsloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if well done) the least important information. But with every step you gain less and need to shave of more important things.
> And also be effected by who did the quantization?
My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.