If you can run this using ollama, then you should be able to use https://www.continue.dev/ with both IntelliJ and VSCode. Haven’t tried this model yet - but overall this plugin works well.
https://github.com/ggerganov/llama.cpp/issues/8519
https://github.com/ggerganov/llama.cpp/issues/7727
Mamba support was added in March of this year:
https://github.com/ggerganov/llama.cpp/pull/5328
I have not yet seen a PR to address Mamba2.
My main complain is that the chat sometimes fails to correctly render some GPT-4o output (e.g. LaTeX expressions), but it's mostly fixed with a custom system prompt. It also significantly reduces the battery life of my Macbook M1, but that's expected.