If this quantization method works with smaller models, it would enable running up to 33B models with only 12GB VRAM.
Especially important for democratizing access to Mistral MoE new model.
Especially important for democratizing access to Mistral MoE new model.