Mixtral 8x7B: A sparse Mixture of Experts language model
arxiv.org
arxiv.org
Something that has come to light since the release of the weights, and not mentioned in this paper is that it looks like fairly likely that the 8 experts were all seeded by Mistral 7B and subsequently diverged. This has generated a lot of experimentation in the local LLM community with cloning models as a way to cheaply generate experts.
It was generally thought likely that training an 8x7B network would be as much work as training 8 7B networks, but this seems not to have been true for Mistral, which is super interesting.
There's still a lot of rapid innovation happening in this space, with papers like Calm from DeepMind this week, and a lot of the adhoc experimental layer combining happening in the wild, (see, e.g. Goliath-120b), I think we're likely to see some pretty interesting architectural improvements this year in the LLM space.
Calm seems to point the way to a next step after MoE, and models like Goliath seem to indicate that even a really really lazy version of Calm (no Linear layer combination, just literally alternating layers at full weights) can be very impactful. Overall I think we will see really, really strong models that are performant on consumer hardware in 2024, likely first half of this year.
So far, the main consumer platform capable of running it without 'ruining' the quality of its output (with high levels of quantization) is the newer Apple Silicon Macs with unified memory - generally >=48GB. It can apparently be done on 32 or 36GB, but there's not much headroom.
Edit: As coder543 points out, yes - you can run it without more lossy levels of quantization on multi-GPU setups providing those have enough combined vram.
Mixtral has been MLXd already? Write ups, if any?
However, there's one curious thing: llama.cpp _always_ leads to GPU throttling on Apple Silicon (e.g. the M1Max GPU will go from 1200MHz to around 700MHz), and then fully saturates it. In the rare cases I could get MLX to stay on the GPU, it was able to keep it at the maximum clock rate. However the unpredictable pauses and seemingly unoptimized prompt processing makes it hard to pick a winner in end-to-end tokens/s
https://simonwillison.net/2023/Dec/18/mistral/
I liked "Using Llamafile’s OpenAI API endpoint" described there, using Justine Tunney's llamafiles for Mixtral, but the article link is out of date, as the models have been replaced with newer: https://huggingface.co/jartine
For the amount of money you're talking about, you could also buy two 3090s (~$750 each on eBay) and have 48GB of VRAM to run with less quantization at full speed.
M-series Macs are surprisingly flexible platforms, but they're not "the only" consumer platform that can do Mixtral.
And you're right, you can run it on a multi-GPU setup if you're so inclined.
My 14" M1 Max does around 30t/s on Mixtral Q4_K_M.
That was my experience as well - 3-bit version is pretty good.
I also tried 2-bit version, which was disappointing.
However, there is a new 2-bit approach in the works[1] (merged yesterday) which performs surprisingly well for Mixtral 8x7B Instruct with 2.10 bits per weight (12.3 GB model size).
After trying the various options for running locally, I have settled on just using Ollama - really convenient and easy, and the serve APIs let me use various LLMs in several different (mostly Lisp) programming languages.
With excellent resources from Hugging Face, tool providers, etc., I hope that the user facing interface for running LLMs is simplified even further: enter your hardware specs and get available models filtered by what runs on a user’s setup. Really, we are close to being there.
Off topic: I hope I don’t sound too lazy, but I am retired (in the last 12 years before retirement I managed a deep learning team at Capital One, worked for a while at Google and three other AI companies) and I only allocate about 2 hours a day to experiment with LLMs so I like to be efficient with my time.
1) He doesn't have an ax to grind / an LLM to pimp out, so he's relatively even-handed
2) He uses the same (secret) test data for each model, so his testing is resistant to cherry-picking/finetuning on tests
3) He likes weirdo role-play prompting, so he has a very good sense of the edges of refusal and alignment tuning
4) He picks up stuff well before it hits the only other fair testing I know of, the chat arena
5) I think asking stuff in German is at worst neutral, and at best useful for testing capacity in edge cases.
Practically speaking, his 'preferred' non-giant models, Nous-Capybara-34B and Mixtral both are excellent in comparison with some of the others he looks at, and good recommendations.
That said, I'd like to see a test suite that GPT-4 fails at, or struggles at, at least. And, it would save him a lot of time if he could get something automated together, it's clearly a lot of effort to hand test all those models.
I agree that better ways to evaluate models would be super super useful, and benchmarks like MMLU and whatever's next will continue to be helpful ("real" science). And, it seems like there may even be some benefits for models to training to 'ace the test' more broadly, which is interesting, and ties to some educational theories in teaching humans.
However, one area that open tests can't excel in is in this "fair" evaluation arena -- and I do think that private tests have some value there, to the extent that they can show utility and maintain trust. I don't make any claims about German sex role-play being a good or bad start for these, though.
For example I'm pretty sure Mistral models are better at french, so doing a benchmark using french only would be advantageous for them.
If you want to compare all models, better use english. Because now his benchmark just show which models is better at german.
That being said, it's still a very welcomed benchmark.
3090s are consumer grade and common on gaming rigs. I’m hoping game devs start experimenting with locally deployed Mixtral in their games. e.g. something like CIV but with each leader powered via LLM
Running LLMs locally to create custom dialogue for games is still years away.
Or even better look at something like phi-2. It's likely to go even lower. I'm sure there are people here who can detail more specifics.
Its 1/24 and we're already there more or less.
I had think about this, you need a small LLM for a game, you do not need it to know about movies, music bands, history, coding and all the text on the internet. I am thinking we need a small model similar to phi-2 , trained only on basic stuff then you would train it on the game world lore. Then the game would also use some "simpler" graphics( we had good looking game graphics decades ago so I do not think you could be limited to text adventures or 2D graphics, just you need some simpler and optimized graphics, it is always interesting when someone shows his unreal demo that uses most RAM/VRAM for a simple demo level then a giant game like GTA5)
What do you mean? Most gamers do have an nvidia GPU.
Edit: unless you talk about mobile gamers, and not PC gamers?
https://store.steampowered.com/hwsurvey/Steam-Hardware-Softw...
You can buy a whole PC for that, I refuse to believe that a GPU priced that highly is "consumer grade" and "common".
Are there any GPUs that are good for LLMs or other genAI that aren't absurdly priced? Or ones specifically designed for AI rather than gaming graphics?
Apple is the one doing the best in terms of making consumer-friendly hardware that can perform AI/ML tasks...but that involves a different problem regarding video games.
https://store.steampowered.com/hwsurvey/Steam-Hardware-Softw...
3090 isn't even in the top 30.
Search the page for 3090 and see for yourself, it's on the list twice.
In the case of high-end video games, that's unlikely.
The bigger problem is memory capacity and bandwidth, but I suspect folks will eventually figure out some sort of QoS setup to let the system crunch LLMs using otherwise unused/idle resources.
An Apple M2 Pro with 32GB of RAM is in the same price range as a gaming PC with a 3090, but its another example of normal people with moderately high performance systems "accidentally" being able to run a GPT-3.5 comparable model.
If you have an Apple meeting these specs and want to play around, LLM Studio is open source and has made it really easy to get started: https://lmstudio.ai/
I hope to see a LOT more hobby hacking as a result of Mixtral and successors.
Llmstudio is, but I suspect that was a typo in their comment. https://github.com/TensorOpsAI/LLMStudio
I tried using Ollama on my machine (same specs as above) and it told me I needed 49gb RAM minimum.
ollama run dolphin-mixtral:8x7b-v2.5-q3_K_SEDIT: I just checked, it runs great, thanks.
Is there any speed/performance/quality/context size/etc. advantage to using LLM Studio or any of the other *llama tools that require more setup than downloading and running a single llamafile executable?
1) Download llamafile[1] (30.03 GB): https://huggingface.co/jartine/Mixtral-8x7B-Instruct-v0.1-ll...
2) chmod +x mixtral-8x7b-instruct-v0.1.Q5_K_M.llamafile
3) ./mixtral-8x7b-instruct-v0.1.Q5_K_M.llamafile
[0] https://hacks.mozilla.org/2023/11/introducing-llamafile/
ollama pull mixtral
For a chatgpt-esk web ui
https://github.com/ollama-webui/ollama-webui
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v ollama-webui:/app/backend/data --name ollama-webui --restart always ghcr.io/ollama-webui/ollama-webui:main
Navigate to http://localhost:3000
You can also use ollama in langchain.
There's also developments and experimentation in making it more factual via dpo and laser(afaik so far not very unsuccessful).
Mixtral of experts - https://news.ycombinator.com/item?id=38598559 - Dec 2023 (300 comments)
Mistral-8x7B-Chat - https://news.ycombinator.com/item?id=38594578 - Dec 2023 (69 comments)
Mistral "Mixtral" 8x7B 32k model [magnet] - https://news.ycombinator.com/item?id=38570537 - Dec 2023 (239 comments)
In particular, Mixtral vastly outperforms Llama 2 70B on mathematics, code generation, and multilingual benchmarks
I'm interested in seeing how it does with mathematics. That has always seemed like a particular weakness that no one yet has effectively cracked.I doubt it will be ever 'cracked' with better LLM's, only multimodal ones with access to program execution and calculators.
FWIW I don't agree with this in a theoretical sense. The reason LLMs can do as much as they can is because next token prediction attempts to infer a world model for the processes that generated each next-token in the training set. I don't see a reason that this would preclude learning arithmetic in order to better predict next tokens that require arithmetic.
I'd guess that arithmetic will become suddenly reliable with one of the next significant (e.g. 2-5x) jumps in parameter count.
Any perceived arithmetic ability is just a textual coincidence.
I agree that a sufficiently intelligent LLM, like 300+ IQ would have an excellent model of how multiplication works. It may even assist in finding new theorems, but a calculator will always be better at 926*725.
There's nothing "textual" about the tokens. They are arbitrary identifiers. There's nothing "textual" about transformers! The fact that e.g. GPT-4 can accept images as input, and that its textual performance improved as a result, and that transformers are also being used for text-to-speech models should have already communicated this.
> 2. Next token prediction is a terrible way to perform arithmetic.
This is just attempting to resolve our disagreement with pure assertion. It's certainly less efficient to use an artificial intelligence to do arithmetic. But whether it's efficient is a different question than how likely it is to be possible.
> 3. Perhaps most importantly, the loss function does not incentivise being good at arithmetic at all.
This is blatantly untrue. The same argument would suggest that LLMs can't do anything that wasn't exactly in their training set already. But they can.
2. it's a strong assertion but it is true. it's kind of inherent in the name; next word prediction. Why would you want a calculator to be making predictions? I agree that it's worth understanding whether it's possible, but your original point was that 'arithmetic will become suddenly reliable' - this is what I am disagreeing with.
3. this links to the above point - you are right: LLM's can already perform arithmetic, this could be easily proved by loading up gpt-2. but again, your original point was that LLM's will 'figure out' arithmetic - I don't believe they will. The loss function does not incentivize being good at arithmetic, it incentivizes making good predictions that are close to the true value. The loss function will not penalize an LLM who predict 'good' is the next word when it should have been 'great'. while it might penalize '2+2=3; since this sequence would be strongly represented in the training set, it's not going to penalize the model getting '1234*5678' one digit off - which is the problem.
> Why would you want a calculator to be making predictions?
If the predictions are correct, why wouldn't I? You are objecting to the entire concept of LLMs at this point, there's nothing specific to arithmetic here.
I'm guessing we'll see llms that do input>program(s)>run>summarize>output
I'm not really disagreeing - I just think llms will do "more of the work" themselves by way of writing and running prolog programs, symbolic math (Julia etc) and running theorem provers.
For example, they never really say how they trained the experts or which dataset they used.
Is this the current standard in the field?
It’s becoming pretty common, yeah. The two things you mentioned: training particulars and dataset mixture are also basically the only competitive advantage companies have. Since the code/architecture is trivial to reproduce, anyone with enough money can make a competing model “easily”.
OpenAI started this trend and cemented it with GPT4’s “technical report” which didn’t even specify the number of parameters in the model. They’ve been historically vague about their dataset for far longer than that though.
The advancement in text only models has been amazing, but a lot of the 'emergent' behavior in GPT-4 may be because of multimodal training and not just MoE or parameter sizes.
I'll be curious to see if multimodal smaller models see similar leaps.
Meta also released a (non-commercial) multimodal model among 6 modalities: https://ai.meta.com/blog/imagebind-six-modalities-binding-ai...
The model weights seem to be under a non-commercial license, not true open source, but it is "open access" as you requested.
It would be nice if someone trained a CogVLM-compatible model from scratch under an open source license.
For CLI, Ollama: https://ollama.ai/library/mixtral
With UI, GPT4All: https://gpt4all.io/index.html (doesn't yet support Mixtral)
In-app, superagent.sh: https://github.com/homanp/superagent
LM Studio is another option.
Given their high quality releases so far they means exciting times for open source LLMs.
It looks like each expert is used interchangeably with no clear pattern. And earlier they say "Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic."
So then, what is the point of the "expert"?
Could this extra performance just be through the 8 expert architectural design, and not be based on the underlying training material? For example, if all 8 experts were ArXiv papers, would the performance be different?
~13B models should work well with plenty of room for other applications. Lately, I’ve heard good things about Solar10B, but new models come in a dozen by day, so it might have already been changed.
I don't see any answer to this in the paper, although I only skimmed it.
Each of these 8 models were 7B models. What about using 80 x tinyllama 1B models?
Manually combining specialist variants is a known technique. This paper automates it with a router component which mixes 2 sub-models at any given time. Training 8 slight variants of a base seems safe and configurable compared to n > 16 specialists. The latter seems like the parts could interact unpredictably.
Also, the memory usage seems predictable: it follows 2^m memory conventions by mixing 2 models at a time, so ~2x the memory is actively used at a time. I'm not up to date on the hardware implications, so it might not mean anything yet. It might one day if this approach works well enough to design around.
I've not been following their releases too well but it seemed they were very much on the side of releasing models asap.
The paper just came out today: https://twitter.com/dchaplot/status/1744547220983005478l
- 'sparse mixture of experts' means each layer contains 8 mini-neural-nets (experts), each trained to be good at certain data/tasks
- a router network passes each token (word) to 2 experts which suit it best
- since only 2/8 experts are utilized, the model effectively only uses 13/47B parameters during inference (text generation)
- this expert mechanism makes it very efficient and effective (uses less params, is able to speciailize to tokens)
- beats llmana 70b and gpt-3.5, especially at math and coding and language
- has fine tuned model able to beat gemini pro, as well as llmana 70b and gpt-3.5