Unquantized model is here: https://huggingface.co/152334H/miqu-1-70b-sf
This strikes me as less a leak and more clever marketing from Mistral.
This strikes me as less a leak and more clever marketing from Mistral.
Clearly we should train a diffusion model to denoise the weights of LLM transformer models. Yo dawg.