Codestral Mamba
mistral.ai
mistral.ai
If they had linked to the instructions in their post (or better yet a link to a one click install of a VS Code Extension), it would help a lot with adoption.
(BTW I consider it malpractice that they are at the top of hacker news with a model that is of great interest to a large portion of the users where and they do not have a monetizable call to action on the page featured.)
https://github.com/ggerganov/llama.cpp/issues/8519
https://github.com/ggerganov/llama.cpp/issues/7727
Mamba support was added in March of this year:
https://github.com/ggerganov/llama.cpp/pull/5328
I have not yet seen a PR to address Mamba2.
My main complain is that the chat sometimes fails to correctly render some GPT-4o output (e.g. LaTeX expressions), but it's mostly fixed with a custom system prompt. It also significantly reduces the battery life of my Macbook M1, but that's expected.
Also, doesn't seem to have a freemium tier...need to start paying even before trying it out ?
"Our API is currently available through La Plateforme. You need to activate payments on your account to enable your API keys."
>Monthly subscription based, free until 1st of August
They just released weights, and being a for profit, need to optimize for making money, not eyeballs. It seems wise to guide people to the API offering.
Github Copilot makes +100M/year, if not way way more.
Having a VS Code extension for Mistral would be a revenue stream if it was one-click and better or cheaper than Github Copilot. It is malpractice in my mind to not be doing this if you are investing in creating coding models.
I assumed they meant free x local. It doesn't seem rational to make this one paid: its significantly smaller than their better model, and even more so than Copilot's.
plus they’ve got some kinda enterprise/team offering; assuming they charge extra there, I could easily see $100M ARR
but that’s pure conjecture, and generous at that; I don’t think we have any hard numbers
Would it help signal competency? They're a small team focused on making models, not VS Code extensions.
Would they do M&A? The founding team is ex-Googlers and has found significant attention in the MBA world via being an EU champion.
Half-baked idea that obviously the models would need to be tuned for different languages / for specific knowledge, therefore countries would pay to do that.
There were many ideas like that, none of them panned out, hence the defenestration. All love for the guy, he did a very, very good thing. It's just meaningless to invoke it here, not only because it's completely off-topic, if anything that's already the play as the EU champion, and because the Stability gentleman was just thinking out loud, nothing more.
If your country believes guns=bad nipples=good war=hell but you get your novels and history books written by an LLM trained by people who believe guns=good nipples=bad war=heroic it would be naive to expect the output to reflect your values and not theirs.
Even close allies of the US would be nervous to have such power in the hands of American multinational corporations alone - so the French state could be very eager for Mistral to produce a competitive product.
The issue with using any local models for code generation comes up with doing so in a professional context: you lose any infrastructure the provider might have for avoiding regurgitation of copyright code, so there's a legal risk there. That might not be a barrier in your context, but in my day-to-day it certainly is.
Thank you for sharing, this is almost exactly what I've been looking for, for ages!
It appears that for the $10/month you just get access to additional features (e.g. bigger context) + a budget of $5/month of credits. The credits possibly translate 1:1 to the usage costs of the underlying models you chose.
I asked it to add some missing documentation with the Claude 3.5 Sonnet model to a medium-sized Python file (2k lines) as a first test and that used up 13 cents of the credits. If it were working in a really productive way (where I'd also include more files as context), I'd probably burn through $20-50/day.
I'm wondering why they are not publicly showing some more transparent pricing. Yeah, this method will bump their signup numbers for investors, but it also absolutely wrecks their churn numbers, when people immediately cancel in the first hour.
It also has a pay as you go chat which is integrated into the IDE w/ hotkeys, however the above subscription gives you $5 free credits for it a month. This doesn't call their own model but stuff like gpt-4o which is presumably why they don't offer unlimited free usage.
There is a separate, optional feature called Supermaven Chat which is basically a wrapper around GPT-4o / Claude 3.5 where you can chat about fragments of your code, refactor them etc. This costs money, but you can also supply your own API key for OpenAI/Anthropic models. Or just use the web interface of any LLM if you just want to chat about your codebase.
- I don't remember the shortcuts most of the time.
- When I run completions I double take and realise they're wrong.
- I am not a good source of data.
All this information is being fed back into the model as positive feedback. So perhaps reason for it to have gone downhill.
I recall it being amazing at coding back in the day, now I can't trust it.
Of course, it's anecdotal which is also problematic in itself but I have definitely noticed the issue where it will fail and stop autocompleting or provide completely irrelevant code.
Pure speculation of course.
function! GetSurroundingLines(n)
let l:current_line = line('.')
let l:start_line = max([1, l:current_line - a:n])
let l:end_line = min([line('$'), l:current_line + a:n])
let l:lines_before = getline(l:start_line, l:current_line - 1)
let l:lines_after = getline(l:current_line + 1, l:end_line)
return [l:lines_before, l:lines_after]
endfunction
function! AIComplete()
let l:n = 256
let [l:lines_before, l:lines_after] = GetSurroundingLines(l:n)
let l:prompt = '<PRE>' . join(l:lines_before, "\n") . ' <SUF>' . join(l:lines_after, "\n") . ' <MID>'
let l:json_data = json_encode({
\ 'model': 'codellama:13b-code-q6_K',
\ 'keep_alive': '30m',
\ 'stream': v:false,
\ 'prompt': l:prompt
\ })
let l:response = system('curl -s -X POST -H "Content-Type: application/json" -d ' . shellescape(l:json_data) . ' http://localhost:11434/api/generate')
let l:completion = json_decode(l:response)['response']
let l:paste_mode = &paste
set paste
execute "normal! a" . l:completion
let &paste = l:paste_mode
endfunction
nnoremap <leader>c :call AIComplete()<CR>All of those are FIM capable, but especially deepseek-v2-lite is very picky with its prompt template so make sure you use it correctly...
Depending on your hardware codestral-22B might be fast enough for everything, but for me it's a bit to slow...
If you can run it deepseek v2 non-light is amazing, but it requires loads of VRAM
EDIT: nvm, my mistake looks like it works fine https://github.com/ollama/ollama/issues/5403
Deepseek uses a 4k sliding window compared to Codestral Mamba's 256k+ tokens
Has anyone tried this? And then, is it fast(er)?
This is what introduced me to them. May be a bit outdated at this point.
The paper author has a blog series but I don't think it's for general public https://tridao.me/blog/2024/mamba2-part1-model/
> We have tested Codestral Mamba on in-context retrieval capabilities up to 256k tokens
Why only 256k tokens? Gemini's context window is 1 million or more and it's (probably) not even using Mamba.
This is coming from someone that understands the general concepts of how LLMs work but only used the general publicly available tools like ChatGPT, Claude, etc.
I want to see if I have any hardware I can stress and run something locally, but don’t know where to start or even what are the available options.
https://github.com/open-webui/open-webui
This will install ollama and open web GUI:
For GPU support run:
docker run -d -p 3000:8080 --gpus=all -v ollama:/root/.ollama -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:ollama
Use for CPU only support:
docker run -d -p 3000:8080 -v ollama:/root/.ollama -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:ollama
https://github.com/oobabooga/text-generation-webui
It's like you hate settings, features, and access to many backends!
Check hardware requirements here: https://rahulschand.github.io/gpu_poor/
You can run a 7b on most modern hardware.How fast will vary.
To run 30-70b models you're getting in the realm of needing 24gb or more of vRAM.
https://www.reddit.com/r/LocalLLaMA/comments/1cj4det/llama_3...
I haven't used any agent frameworks other than messing around with langchain a bit so I can't speak to how that would effect things.
I can't agree with "very bad". Maybe your standards are set by the best, largest models, but have a little perspective: a modern 7b model is a friggin magical piece of software. Fully in the realm of sci-fi until basically last Tuesday. It can reliably summarize documents, bash a 30 minute rambling voice note into a terse proposal, and give you social counseling at least on par with r/Relationship_Advice. It might not always get facts exactly right but it is smart in a way that computers have never been before. And for all this capability, you can get it running on a computer a decade old, maybe even a Raspberry Pi or a smartphone.
To answer the parent: Download a "gguf" file (blob of weights) of a popular model like Mistral from HugginFace. Git pull and compile llama.cpp. Run ./main -m path/to/gguf -p "prompt"
Not sure what the equivalent is for image generation; it's either https://www.reddit.com/r/StableDiffusion/ or one of the related subreddits it links to.
Sadly, I've yet to find anyone doing "daily ML-hobbyist news" content creation, summarizing the types of articles that appear on these subreddits. (Which is a surprise to me, as it's really easy to find e.g. "daily homelab news" content creators. Please, someone, start a "daily ML-hobbyist news" blog/channel! Given that the target audience would essentially be "people who will get an itch to buy a better GPU soon", the CPM you'd earn on ad impressions would be really high...)
---
That being said, just to get you started, here's a few things to know at present about "what you can run locally":
1. Most models (of the architectures people care about today) will probably fit on a GPU which has something like 1.5x the VRAM of the model's parameter-weights size. So e.g. a "7B" (7 billion parameter-weights) model, will fit on a GPU that has 12GB of VRAM. (You can potentially squeeze even tighter if you have a machine with integrated graphics + dedicated GPU, and you're using the integrated graphics as graphics, leaving the GPU's VRAM free to only hold the model.)
2. There are models that come in all sorts of sizes. Many open-source ML models are huge (70B, 120B, 144B — things you'd need datacenter-class GPUs to run), but then versions of these same models get released which have been heavily cut down (pruned and/or quantized), to force them to fit into smaller VRAM sizes. There are 5B, 3B, 1B, even 0.5B models (although the last two are usually special-purpose models.)
3. Surprisingly, depending on your use-case, smaller models (or small quants of larger models) can "mostly" work perfectly well! They just have more edge-cases where something will send them off the rails spiralling into nonsense — so they're less reliable than their larger cousins. You might have to give them more prompting, and try regenerating their output from the same prompt several times, to get good results.
4. Apple Silicon Macs have a GPU and TPU that read from/write to the same unified memory that the CPU does. While this makes these devices slower for inference than "real" GPUs with dedicated VRAM, it means that if you happen to own a Mac with 16GB of RAM, then you own something that can run 7B models. AS Macs are, oddly enough, the "cheapest" things you can buy in terms of model-capacity-per-dollar. (Unlike a "real" GPU, they won't be especially quick and won't have any capacity for concurrent model inference, so you'd never use one as a server backing an Inference-as-a-Service business. But for home use? No real downsides.)
After ChatGPT released, there was a lot of hype in the space but open source was far behind. Iirc the best open foundation LLM that existed was GPT-2 but it was two generations behind.
Awhile later Meta released LLaMA[1], a well trained base foundation model, which brought an explosion to open source. It was soon implemented in the Hugging Face Transformers library[2] and the weights were spread across the Hugging Face website for anyone to use.
At first, it was difficult to run locally. Few developers had the system or money to run. It required too much RAM and iirc Meta's original implementation didn't support running on the CPU but developers soon came up with methods to make it smaller via quantization. The biggest project for this was Llama.cpp[3] which probably is still the biggest open source project today for running LLMs locally. Hugging Face Transformers also added quantization support through bitsandbytes[4].
Over the next months there was rapid development in open source. Quantization techniques improved which meant LLaMA was able to run with less and less RAM with greater and greater accuracy on more and more systems. Tools came out that were capable of finetuning LLaMA and there were hundreds of LLaMA finetunes that came out which finetuned LLaMA on instruction following, RLHF, and chat datasets which drastically increased accuracy even further. During this time, Stanford's Alpaca, Lmsys's Vicuna, Microsoft's Wizard, 01ai's Yi, Mistral, and a few others made their way onto the open LLM scene with some very good LLaMA finetunes.
A new inference engine (software for running LLMs like Llama.cpp, Transformers, etc) called vLLM[5] came out which was capable of running LLMs in a more efficient way than was previously possible in open source. Soon it would even get good AMD support, making it possible for those with AMD GPUs to run open LLMs locally and with relative efficiency.
Then Meta released Llama 2[6]. Llama 2 was by far the best open LLM for its time. Released with RLHF instruction finetunes for chat and with human evaluation data that put its open LLM leadership beyond doubt. Existing tools like Llama.cpp and Hugging Face Transformers quickly added support and users had access to the best LLM open source had to offer.
At this point in time, despite all the advancements, it was still difficult to run LLMs. Llama.cpp and Transformers were great engines for running LLMs but the setup process was difficult and required a lot of time. You had to find the best LLM, quantize it in the best way for your computer (or figure out how to identify and download one from Hugging Face), setup whatever engine you wanted, figure out how to use your quantized LLM with the engine, fix any bugs you made along the way, and finally figure out how to prompt your specific LLM in a chat-like format.
However, tools started coming out to make this process significantly easier. The first one of these that I remember was GPT4All[7]. GPT4All was a wrapper around Llama.cpp which made it easy to install, easy to select the LLM that you want (pre-quantized options for easy download from a download manager), and a chat UI which made LLMs easy to use. This significantly reduced the barrier to entry for those who were interested in using LLMs.
The second project that I remember was Ollama[8]. Also a wrapper around Llama.cpp, Ollama gave most of what GPT4All had to offer but in an even simpler way. Today, I believe Ollama is bigger than GPT4All although I think it's missing some of the higher-level features of GPT4All.
Another important tool that came out during this time is called Exllama[9]. Exllama is an inference engine with a focus on modern consumer Nvidia GPUs and advanced quantization support based on GPTQ. It is probably the best inference engine for squeezing performance out of consumer Nvidia GPUs.
Months later, Nvidia came out with another new inference engine called TensorRT-LLM[10]. TensorRT-LLM is capable of running most LLMs and does so with extreme efficiency. It is the most efficient open source inferencing engine that exists for Nvidia GPUs. However, it also has the most difficult setup process of any inference engine and is made primarily for production use cases and Nvidia AI GPUs so don't expect it to work on your personal computer.
With the rumors of GPT-4 being a Mixture of Experts LLM, research breakthroughs in MoE, and some small MoE LLMs coming out, interest in MoE LLMs was at an all-time high. The company Mistral had proven itself in the past with very impressive LLaMA finetunes, capitalized on this interest by releasing Mixtral 8x7b[11]. The best accuracy for its size LLM that the local LLM community had seen to date. Eventually MoE support was added to all inference engines and it was a very popular mid-to-large sized LLM.
Cohere released their own LLM as well called Command R+[12] built specifically for RAG-related tasks with a context length of 128k. It's quite large and doesn't have notable performance on many metrics, but it has some interesting RAG features no other LLM has.
More recently, Llama 3[13] was released which similar to previous Llama releases, blew every other open LLM out of the water. The smallest version of Llama 3 (Llama 3 8b) has the greatest accuracy for its size of any other open LLM and the largest version of Llama 3 released so far (Llama 3 70b) beats every other open LLM on almost every metric.
Less than a month ago, Google released Gemma 2[14], the largest of which, performs very well under human evaluation despite being less than half the size of Llama 3 70b, but performs only decently on automated benchmarks.
If you're looking for a tool to get started running LLMs locally, I'd go with either Ollama or GPT4All. They make the process about as painless as possible. I believe GPT4All has more features like using your local documents for RAG, but you can also use something like Open WebUI[15] with Ollama to get the same functionality.
If you want to get into the weeds a bit and extract some more performance out of your machine, I'd go with using Llama.cpp, Exllama, or vLLM depending upon your system. If you have a normal, consumer Nvidia GPU, I'd go with Exllama. If you have an AMD GPU that supports ROCm 5.7 or 6.0, I'd go with vLLM. For anything else, including just running it on your CPU or M-series Mac, I'd go with Llama.cpp. TensorRT-LLM only makes sense if you have an AI Nvidia GPU like the A100, V100, A10, H100, etc.
[1] https://ai.meta.com/blog/large-language-model-llama-meta-ai/
[2] https://github.com/huggingface/transformers
[3] https://github.com/ggerganov/llama.cpp
[4] https://github.com/bitsandbytes-foundation/bitsandbytes
[5] https://github.com/vllm-project/vllm
[6] https://ai.meta.com/blog/llama-2/
[7] https://www.nomic.ai/gpt4all
[9] https://github.com/turboderp/exllamav2
[10] https://github.com/NVIDIA/TensorRT-LLM
[11] https://mistral.ai/news/mixtral-of-experts/
[12] https://cohere.com/blog/command-r-plus-microsoft-azure
[13] https://ai.meta.com/blog/meta-llama-3/
[14] https://blog.google/technology/developers/google-gemma-2/
I'm pretty sure CodeLlama is out of date now. I've heard DeepSeek LLMs are good and DeepSeek-Coder-V2-Instruct was released recently. With the good reputation and its massive size (236b) I'd guess it is the best coding LLM, but if it's not being trained efficiently, maybe Codestral and Codestral Mamba come close.
I don't think the best coding LLMs are close to GitHub Copilot but I could be wrong since I'm just relaying information that I've heard secondhand.
[1] https://ai.meta.com/blog/code-llama-large-language-model-cod...
[2] https://mistral.ai/news/codestral/
[3] https://github.com/deepseek-ai/DeepSeek-Coder-V2
[4] https://developers.googleblog.com/en/gemma-family-expands-wi...
I've personally went back to the browser with Claude 3.5 Sonnet (and the projects + artifacts feature), as it is one of the most industrious ones, and I really like the UX of artifacts + it integrates new code well into existing code you paste into it.
In the end I think it also often comes down to what languages/frameworks you are using and how well the LLM/product handles it, so I'd still recommend to test around. E.g. some of the main frameworks I'm working with on a daily basis went through big refactors/interface changes 1-2 years ago, and I stopped using ChatGPT because it had a strong tendency to produce code based on the old interfaces/paradigms.
Aider[0] is also quite interesting, especially when it comes to more significant refactorings in the codebase and has gotten quite good with that with the last few bigger model releases, but it takes same time to get used to and doesn't have good IDE-integration.
> Awhile later Meta released LLaMA[1],
I think Stable Diffusion was first to release a SOTA model (August 2022) that worked locally, not in language but image generation, but it set the tone for Meta. LLaMA only came in February 2023.
> The company Mistral had proven itself in the past with very impressive LLaMA finetunes
Mistal is not a finetune of LLaMA, it is a model trained from scratch. Also, Mistral was most of the time better than LLaMA during this period.
> Quantization techniques improved which meant LLaMA was able to run with less and less RAM with greater and greater accuracy
Quantization does not improve accuracy, except if you trade off precision for longer context maybe, but not on similar prompts. It is like JPEG compression, the original is always better for a specific image, but for the same byte size you get more resolution from JPEG than say... a PNG.
Sure, I was only covering LLMs though. If I wanted to cover image generation models and tools as well, the comment would be double its size.
> Mistal is not a finetune of LLaMA, it is a model trained from scratch. Also, Mistral was most of the time better than LLaMA during this period.
Oh, that's right. Iirc it was just the Llama 2 architecture that was used with sliding window attention.
> Quantization does not improve accuracy, except if you trade off precision for longer context maybe, but not on similar prompts. It is like JPEG compression, the original is always better for a specific image, but for the same byte size you get more resolution from JPEG than say... a PNG.
I'm well aware of how quantization works. I meant quantization methods were increasingly able to retain accuracy. Such as methods which quantize less important weights more heavily, improving accuracy for the same LLM size.
I haven't seen any good non-paper explainers yet.
Also they did the thing that junior developers tend to do, where you have a race condition of some sort, and they just work around it by adding some if checks. The app is at around 400 lines right now, it works but feels pretty brittle. Adding a tiny feature here or there breaks something else, and GPT does the wrong thing half the time.
All in all, I'm not complaining, because I made an app in two days, but it won't replace a developer yet, no matter how much I want it to.
For example, I'm currently working on a Rust/Qt desktop app so I have a project with the whole Qt6 book attached to ask questions about Qt, a project with my SQL schema and ORM/Sqlite docs to ask questions about the app's data and generate models without dealing with hallucinations, a project with all my QML files and Rust QML element code, a project with a bunch of Rust crate docs, and so on and on.
GPTs allow attaching files too but Claude Projects dump the entire contents of the files into the context rather than trying to do some hacky RAG that never works like I want it to.
The UX of drag and dropping a few monolithic markdown files to include entire chunks of a large project outweighs the downsides of including irrelevant context in my experience.
The more context you give the llm, the better.
The key takeaway from that paper is to keep your instructions/questions/direction in the beginning or at the end of the context. Any information can go anywhere.
Not to be too dismissive, it's a good paper, but we're one year further and in practice this issue seems to have been tackled by training on better data.
This can differ a lot depending on what model you're using, but in the case of claude sonnet 3.5, more relevant context is generally better for anything except for speed.
It does remain true that you need to keep your most important instructions at the beginning or at the end however.
c.f.
https://pbs.twimg.com/media/GH2NJMxbYAAcRL3?format=jpg&name=...
I do wish they compare it to codegeex4-all-9b
You just need a Mistral API key: https://console.mistral.ai/api-keys/
> As a tribute to Cleopatra, whose glorious destiny ended in tragic snake circumstances
but according to Wikipedia this is not true:
> When Cleopatra learned that Octavian planned to bring her to his Roman triumphal procession, she killed herself by poisoning, contrary to the popular belief that she was bitten by an asp.
> [A]ccording to the Roman-era writers Strabo, Plutarch, and Cassius Dio, Cleopatra poisoned herself using either a toxic ointment or by introducing the poison with a sharp implement such as a hairpin. Modern scholars debate the validity of ancient reports involving snakebites as the cause of death and whether she was murdered. Some academics hypothesize that her Roman political rival Octavian forced her to kill herself in a manner of her choosing. The location of Cleopatra's tomb is unknown. It was recorded that Octavian allowed for her and her husband, the Roman politician and general Mark Antony, who stabbed himself with a sword, to be buried together properly.
I think this rounds to “nobody really knows.”
The “glorious destiny” seems kind of shaky, too. It’s just a throwaway line anyway.
They're certainly not the first to use Cleopatra this way nor the most egregious, but there are plenty of other random mamba jokes that could have filled in there and both made more sense and been less crass.
Ever heard of Patrice Lumumba? He was a congolese politician involved in its independence and democratization. He was shot with the involvement of western forces. There’s a drink named after him: Hot chocolate with a shot of rum.