This isn’t a fine tune. It’s not even a transformer like Qwen, Llama, GPT, Gemini, or Mistral. It’s another approach to language models based on RNNs
Doesn't the point remain though -- when/where do we demo this?
There is Mamba 7B pre trained model on hugging face/ LM studio right now.
Do you mean include it with the Mixtral "Mixture of Experts" model? I'm not sure Mistral Mamba makes sense, since it's a completely different architecture.