Mixtral 8x22B
mistral.ai
mistral.ai
Output: https://gist.github.com/IAmStoxe/7fb224225ff13b1902b6d172467...
Within the first paragraph, it outputs:
> GET AN ESSAY WRITTEN FOR YOU FROM AS LOW AS $13/PAGE
Thought that was hilarious.
ollama run mixtral:8x22b
EDIT: I like how you ninja-editted your comment ;)
And even the direct tag page: https://ollama.com/library/mixtral:8x22b shows 40-something minutes ago: https://imgur.com/a/WNhv70B
Mixtral-8x22B-v0.1 was released a couple days ago. The "mixtral:8x22b" tag on ollama currently refers to it, so it's what you got when you did "ollama run mixtral:8x22b". It's a base model only capable of text completion, not any other tasks, which is why you got a terrible result when you gave it instructions.
Mixtral-8x22B-Instruct-v0.1 is an instruction-following model based on Mixtral-8x22B-v0.1. It was released two hours ago and it's what this post is about.
(The last updated 44 minutes ago refers to the entire "mixtral" collection.)
ollama run mixtral:8x22b
Error: exception create_tensor: tensor 'blk.0.ffn_gate.0.weight' not found
Update: mixtral:8x22b now points to the instruct model:
ollama pull mixtral:8x22b
ollama run mixtral:8x22bI've long thought that if you want reproducibility and reliability, you need to pin your deps.
So, IMO, the change is very much worth it to reduce confusion going forward.
It typically won't behave like human experts - you might find one of the networks is an expert in determining where to place capital letters or full stops for example.
MoE's do not really improve accuracy - instead they are to reduce the amount of compute required. And, assuming you have a fixed compute budget, that in turn might mean you can make the model bigger to get better accuracy.
During training a routing network is punished if it does not evenly distribute training tokens to the correct experts. This prevents any one or two networks from becoming the primary networks.
The result of this is that each token has essentially even probability of being routed to one of the sub models, with the underlying logic of why that model is an expert for that token being beyond our understanding or description.
Basically you're sharding the neural network by "something" that is itself tuned during the learning process.
Maybe a real expert can confirm if this is correct :)
Edit: Apparently each part of the network is on a separate device. Fascinating! That would also explain why the routing network is trained to choose equally between experts.
I imagine that may reduce quality somewhat though? By forcing it to distribute problems equally across all of them, whereas in reality you'd expect task type to conform to the pareto distribution.
Computational costs, yes. You still take the same amount of time for processing the prompt, but each token created through inference costs less computationally than if you were running it through _all_ layers.
We can't really tell what the router does. There have been experiments where the router in the early blocks was compromised, and quality only suffered moderately. In later layers, as the embeddings pick up more semantic information, it matters more and might approach our naive understanding of the term "expert".
Human brains seem to do something similar, inasmuch as blood flow (and hence energy use) per region varies depending on the current problem.
Or would that make each expert too small to be useful? TinyLlama is 1B and it's almost useful! I guess 8x1B would be Mixture of TinyLLaMAs...
(Notably, MoE should not be conflated with ensemble techniques, which is where you would train entire separate networks, then use heuristic techniques to run inference across all of them simultaneously and combine the results.)
In a longer, but still rough, form...
How these transformers work is roughly:
``` x_{l+1} = mlp_l(attention_l(x_l)) ```
where `x_l` is the hidden representation at layer l, `attention_l` is the attention sublayer at layer l, and `mlp_l` is the multilayer perceptron at sublayer l.
This MLP layer is very expensive because it is fully connected (i.e. every input has a weight to every output). So! MoEs instead of creating an even bigger, more expensive MLP to get more capability, they create K MLP sublayers (the "experts") and a router that decides which MLP sublayers to use. This router spits out an importance score for each MLP "expert" and then you choose the top T MLPs and do an average weighed on importance, so roughly:
``` x_{l+1} = \sum_e mlp_{l,e}(attention_l(x_l)) * importance_score_{l, e} ```
where the `importance_score_{l, e}` is the score computed by the router at layer l for "expert" e. That is, `importance_score_{l} = attention_l(x_l)`. Note that here we are adding all experts, but in reality we choose the top T, often 2, and use that.
[0] some architectures do, in fact, combine domain experts to make a greater whole, but not the currently popular flavor
> MLPs are universal function approximators, but these models are big enough that it is better to train many small functions rather than a single unified function. MoE is a mechanism to force different parts of the model to learn distinct functions.
During inference, instead of a single MLP in each transformer layer, MoEs have `n` MLPs and a single layer "gate" in each transformer layer. In the forward pass, softmax of the gate's output is used to pick the top `k` (where k is < n) MLPs to use. The relevant code snippet in the HF transformers implementation is very readable IMO, and only about 40 lines.
https://github.com/huggingface/transformers/blob/main/src/tr...
These models are actually a collection of weights for different parts of the system. It’s not “one” neural network. Transformers are composed of layers of transformations to the input, and each step can have its own set of weights. There was a recent video on the front page that had a good introduction to this. There is the MLP, there are the attention heads, etc.
With that in mind, a MoE model is basically where one of those layers has X different versions of the weights, and then an added layer (another neural network with its own weights) that picks the version of “expert” weights to use.
Maybe this limit will become a joke when looking back? Can you imagine reaching a trillion tokens context window in the future, as Sam speculated on Lex's podcast?
If we have cheap billion token context windows, 99% of your use cases aren't going to hit anywhere close to that limit and as a result, your models will "just run"
https://arxiv.org/html/2404.08801v1 Meta Megalodon
https://arxiv.org/html/2404.07143v1 Google Infini-Attention
https://arxiv.org/html/2402.13753v1 LongRoPE
and a ton more
A better phrasing would be that they don't allow you to output more than 4k tokens per message.
Same with Anthropic and Claude, sadly.
For something like Mixtral 8X22B with 40B active params you'd looking at the $10M range, and if something gets screwed up during training you can be left with a dud and nothing to show for it, like LLama-2-33B. It's like buying millions worth of lootboxes and hoping something good drops.
During a court case, the other side can demand discovery over your training dataset, for example to see if it contains a particular copyrighted work.
But if you've already deleted the dataset, you're far more likely to win any case against you that hinges on what was in the dataset if the plaintiff can't even prove their work was included.
And you can argue that the dataset was very expensive to store (which is true), and therefore deleted shortly after training was complete. You have no obligation to keep something for the benefit of potential future plaintiffs you aren't even aware of yet.
I've tried LMStudio, but I'm not a fan of the interface compared to OpenAI's. The lack of automatic regeneration every time I edit my input, like on ChatGPT, is quite frustrating. I also gave Ollama a shot, but using the CLI is less convenient.
Ideally, I'd like something that allows me to edit my settings quite granularly, similar to what I can do in OpenLM, with the QoL from the hosted online platforms, particularly the ease of editing my prompts that I use extensively.
Open WebUI is functionally identical to the ChatGPT interface. You can even use it with the OpenAI APIs to have your own pay per use GPT 4. I did this.
https://docs.openwebui.com/getting-started/
IIUC the load balancing page is for people who want to run openwebui at a larger scale
Most other benchmarks don't clearly show the difference between the top models and the rest. This may be because they are older and have been over-optimized or perhaps because they are just easier.
Wish I had invested in the extra 32GB for my mac laptop.
Edit: I haven't owned a laptop for years, probably could have surmised they'd be more user hostile nowadays.
It's complete garbage. And most of the other vendors just copy Apple so even things like Lenovo have the same problems.
The current state of laptops is such trash
People need to vote with their wallet, and not buy stuff that goes against their principles.
> SO-DIMM memory is inherently slower than soldered memory. Moreover, considering the fact that SO-DIMM has a maximum speed of 6,400MHz means that it won’t be able to handle the DDR6 standard, which is already in the works.
The question remains unanswered. Perhaps I didn't see it for sale or Bob in accounting just got one and I didn't want to look like I was copying Bob.
Even at scale this doesn't work. Let's say Lenovo switches to making all of their laptops hot pink with bedazzled rhinestone butterflies and sales plummet. You could argue it was the wrong pink or that the butterflies didn't shimmer enough ... any hypothesis you wish.
The market provides an extremely low information poor signal that really doesn't suggest any course of action.
If we really want something better, there needs to be more fruitful and meaningful communication lines. I've come up with various ideas over the years but haven't really implemented them.
Weird conspiracy theories aside, the low power variant of RAM (LPDDR) has to be soldered onto the motherboard, so laptops designed for longer battery life have been using it for years now.
The good news is that a newer variant of low power RAM has just been standardized that features low power RAM in memory modules, although they attach with screws and not clips.
I don't think it's so much "copying Apple" as much as it's learning that they can do things like solder in the RAM to cut costs and discovering that most people won't care and will still buy it.
[...insert sound of old timers laughing at newbies who declared T480/T490 overhyped for no reason because "no one care about upgradability anymore"...]
:)
Both of which seem to be exactly what GP took issue with, and was calling out.
- We first struggled with limited context windows [solved]
- We had issues with consistent JSON ouput [solved]
- We had rate limiting and performance issues for the large 3rd party models [solved]
- Hosting our own OSS models for small and medium complex tasks was a pain [solved]
Obivously every startup still needs to build up defensibility and focus on differentiating with everything “non-AI”.
Or you can continue selling shovels. Still lots of expensive labeling services out there, to stay in the image-recognition parallel
GPT-4 is SOTA at OCR and sentiment classification, for example.
Would you happen to know of any cool OSS model projects that might be good inspiration for a side project?
Wondering what most people use these local models for
It would flat out embarass alexa. Imagine 'Hal play a movie', or 'Hal play some music' and it's all running locally, with your content.
I think the real value in using local models is exposing them to personal/unique information that only you have, thus getting novel and unique outcomes that no public model could provide.
1. Project 1 - Self Knowledge - Download/extract all of my emails and populate into a vector database, like Chroma[0] - For each prompt do a search of the vector store and return N number of matches - Provide both prompt and search result to LLM, instructing it to use the search result as context or in the answer itself.
2. Project 2 - Chat with a Friend - I exported the chat and text history between me and a good friend that passed away - I created a vector store of our chat history in chunks, each consisting of 6 back-and-forth interactions - When I "chat" with the LLM the a search is first conducted for matching chunks from the vector store and then using those as "style" and knowledge context for a response. Optional: You can use SillyTavern[1] for a more "rich" chat experience
The above lets me chat, at least superficially, with my friend. It's nice for simple interactions and banter; I've found it to be a positive and reflective experience.
u rock
It's coming along but very hit or miss with the small models I'm using (all my poor 6750XT 12GB can manage).
I could be doing something wrong though, even with GPT-4 im struggling a bit - it's very lazy and doesn't want to fully populate the data structure fields (the game file objects can haze dozens/hundredss of fields). I'm probably just not using the right incantation/magic phrase/voodoo?
Write a Golang func, which accepts the path into a .gpx file and outputs a JSON string with points(x=tolal distance in km, y=elevation). Don't use any library.
It says the JSON output is constrained via their platform (on la Plateforme).
Does that mean JSON output is only available in the hosted version? Are there any small models that can be self hosted that output valid JSON.
I would assume so. They probably constrain JSON output so that the JSON response doesn't bork the front-end/back-end of la Plateforme itself as it moves through their code back to you.
Yes, for example this one is optimized for function calling and JSON output: https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B
and so do your competitor's products.
folks like ESR were weak with their definition of open source and it wasn't as good.
stallman said:
... the obvious meaning for the expression “open source software” is “You can look at the source code.” This is a much weaker criterion than free software ...
https://www.gnu.org/philosophy/free-software-for-freedom.htm...
https://www.gnu.org/philosophy/open-source-misses-the-point....
My use case is that I'm more productive working with a LLM but being online is a constant temptation and distraction.
Most of the time I'll reach for offline docs to verify. So the LLM just points me in the right direction.
I also miss Google offline, so I'm working on a search engine. I thought I could skip crawling by just downloading common crawl, but unfortnately it's enormous and mostly junk or unsuitable for my needs. So my next project is how to data-mine common crawl to extract just the interesting (to me) bits...
When I have a search engine and a LLM I'll be able to run my own Phind, which will be really cool.
It has been a huge force multiplier for me and most importantly of all, it removes the dread of not knowing where to start and the dread of sending your inner monologue to someone's stupid cloud.
If you're curious: https://github.com/noman-land/transcript.fish/ though this doesn't include any Mixtral stuff because I don't use it programmatically (yet). I soon hope to use it to answer questions about the episodes like who the special guest is and whatnot, which is something I do manually right now.
pretty sure you can run it un-censored... that would be my use case
Around 5 years ago, it took a lambda user some pretty significant hardware, software and time (around a full night), to try to create a short deepfake. Now, you don't need any fancy hardware and you can have some decent results within 5 min on your average computer.
https://www.nytimes.com/2024/04/08/technology/deepfake-ai-nu...
Mixtral8x22 beats CommandR+, which is at GPT-4-level in LMSYS' leaderboard.
Most people would not consider Command R+ to count as the "best permissively licensed model" since CC-BY-NC is not usually considered "permissively licensed" – the "NC" part means "non-commercial use only"
But because it only activates one expert at a time, it can run on a fast CPU in reasonable time. So 96GB of DDR4 will do. 96GB of DDR5 is better.
It's nice they're using some of the money from their commercial and proprietary models, to improve the state of the art for open source (open weights) models.
How did they weaken their commitment to open weights?
Well, yeah, it's very welcome, but 'history of humanity' is hyperbole given ChatGPT isn't even two years old.
> How did they weaken their commitment to open weights?
Before https://web.archive.org/web/20240225001133/https://mistral.a... versus after https://web.archive.org/web/20240227025408/https://mistral.a... the Microsoft partnership announcement:
> Committing to open models.
to
> That is why we started our journey by releasing the world’s most capable open-weights models
There were similar changes on their about the Company page.
Found it: https://mistral.ai/technology/#pricing
It'd useful to add a link to the blog post. While it's an open model, most will only be able to use it via the API.
I wouldn’t call this runnable on most laptops.
~$8k for an LLM server with 128GB of VRAM vs like $250k+ for 8 H100s.
Keep on doing this great work.
Edit: been using the previous version, seems like this one is even better?
I asked it what it's knowledge cutoff was, and it said 2021-09.
Anyone know why it's trained on such old data?
It turns interpreting the results into an exercise in detecting which models and benchmarks were omitted.
The different teams are learning from each other and pushing boundaries; there's virtually no reason for any of the teams to release a model or product that is somehow inferior to a prior one (unless it had some secondary attribute such as requiring lower end hardware).
We're simply not seeing the ones that came up short; we don't even see the ones where it fell short of current benchmarks because they're not worth releasing to the public.
This is completely normal, the opposite would be strange.
It may even be the case that in measuring against the benchmarks, these product teams sacrifice some real world performance (just as a student that only studies for the SAT might sacrifice some real world skills).
If you compare LLMs by asking them to tell you how to catch dragonflies - the free text chat answer you get will be impossible to objectively evaluate.
Whereas if you propose four ways to catch dragonflies and ask each model to choose option A, B, C or D (or check the relative probability the model assigns to those four output logits) the result is easy to objectively evaluate - you just check if it chose the one right answer.
Hence a lot of the most famous benchmarks are multiple-choice questions - even though 99.9% of LLM usage doesn't involve answering multiple-choice questions.