Mixtral of experts
mistral.ai
mistral.ai
Official post on Mixtral 8x7B: https://mistral.ai/news/mixtral-of-experts/
Official PR into vLLM shows the inference code: https://github.com/vllm-project/vllm/commit/b5f882cc98e2c9c6...
New HuggingFace explainer on MoE very nice: https://huggingface.co/blog/moe
In naive decoding, performance of a bit above 70B (Llama 2), at inference speed of ~12.9B dense model (out of total 46.7B params).
Notes: - Glad they refer to it as "open weights" release instead of "open source", which would imo require the training code, dataset and docs. - "8x7B" name is a bit misleading because it is not all 7B params that are being 8x'd, only the FeedForward blocks in the Transformer are 8x'd, everything else stays the same. Hence also why total number of params is not 56B but only 46.7B. - More confusion I see is around expert choice, note that each token and also each layer selects 2 different experts (out of 8). - Mistral-medium
Source: https://twitter.com/karpathy/status/1734251375163511203
It seems recently OpenAI is the least open startup. Even Gemini talks more about their architecture.
OpenAI still doesn’t openly mention GPT4 is a mixture of experts model.
Already available from both Mistralai and TheBloke https://huggingface.co/mistralai/Mixtral-8x7B-v0.1 https://huggingface.co/TheBloke/Mixtral-8x7B-v0.1-GGUF
The GGUF handling for Mistral's mixture of experts hasn't been finalized yet. TheBloke and ggerganov and friends are still figuring out what works best.
The Q5_K_M gguf model is about 32GB. That's not going to fit into any consumer grade GPU, but it should be possible to run on a reasonably powerful workstation or gaming rig. Maybe not fast enough to be useful for everyday productivity, but it should run well enough to get a sense of what's possible. Sort of a glimpse into the future.
Also, given the insane cost premium apple charges per extra GB of RAM (at least when I was last shopping for a device), do you come out ahead?
Intel Core i9-13900F memory bandwidth: 89.6 GB/s, memory size up to 192 GB
Apple M3 Pro memory bandwidth: 150GB/s, memory size up to 36GB
Apple M3 Max memory bandwidth: 300GB/s, memory size up to 128GB
GeForce RTX 4090 memory bandwidth: 1008 GB/s, memory size 24GB fixed, no more than two cards per PC.
A GPU connected to a PCIe 3.0 x16 electrical uplink would be constrained to ~16GB/s, or ~32GB/s if it were a PCIe 4.0 uplink instead. Although those numbers imply slower bandwidth than CPU inference, that bottleneck would only be when paging in or out (or directly accessing?) layers overflowed to the shared system ram, so they don't really represent much on their own.
> no more than two cards per PC
I've seen quad 4090 builds, e.g. here[0]. What do you mean no more than two cards? Yes, power is definitely an issue with multiple 4090s, though you can limit the max power using `nvidia-smi`, which IME doesn't hurt (mem-bottlenecked) inference.
[0] https://old.reddit.com/r/watercooling/comments/16ed8fu/quad_...
I wish there were affordable platforms with quad DDR5.
I can only speculate that it would help mitigate latency with loose timings on a fast OC among other things.
For me it's OK though, since I want faster compile times anyway, so it's worth the money. To me local LLMs are just a curiosity.
edit: Interesting information here. https://old.reddit.com/r/LocalLLaMA/comments/14ilo0t/extensi...
> RAM speed does not matter. The processing time is identical with DDR-6000 and DDR-4000 RAM.
You'd really expect DDR5-6000 to be advantageous. I think that AMD Ryzen 7xxx can at least take advantage to up to 5600. Does it perhaps not wind up bottlenecking on memory? Maybe quantization plays a role...
That's referring specifically to prompt processing, which uses a batch processing optimization not used in normal inference. The processed prompt can also be cached so you only need to process it again if you change it. Normal inference benefits from faster RAM.
When you say 'token' is this a word? A character? I've never gotten a good definition for it beyond 'a unit of text the llm processes'
Everyone uses (byte pair encoding)[https://en.wikipedia.org/wiki/Byte_pair_encoding] to generate their tokens; the tokens are whatever emerge from this. They will typically correspond to the most common substrings in the training corpus in, handwaving a bit, a max-cover sense; it's an encoding which attempts to best compress the data the tokenizer was trained on.
"Mixtral is a sparse mixture-of-experts network. It is a decoder-only model where the feedforward block picks from a set of 8 distinct groups of parameters. At every layer, for every token, a router network chooses two of these groups (the “experts”) to process the token and combine their output additively.
This technique increases the number of parameters of a model while controlling cost and latency, as the model only uses a fraction of the total set of parameters per token. Concretely, Mixtral has 46.7B total parameters but only uses 12.9B parameters per token. It, therefore, processes input and generates output at the same speed and for the same cost as a 12.9B model."So 45 billion parameters is what they consider their "small" model? I'm excited to see what/if their larger models will be.
According to Wikipedia: Rumors claim that GPT-4 has 1.76 trillion parameters, which was first estimated by the speed it was running and by George Hotz. [1]
[1] https://the-decoder.com/gpt-4-architecture-datasets-costs-an...
Also, if we have been eating up posted "benchmarks" with no way to independently validate them and watching heavily edited video presentations, why can't we trust our wonder kid?
Which also means you can fit the 8 models in a much smaller amount of memory than a 45B model. Latency will also be much smaller than a 45B model, since the next token is always only created by 2 of the 8 models (which 2 models are run is chosen by a different, even smaller/faster, model).
No, Mixture-of-Experts is not stacking finetunes of the same base model.
Made sense to mee on first sight to me, because you don't need to train stuff like syntax and grammar 8 times in 8 different ways.
Also would explain why interference of two 7B models has the cost of running a 12B model.
"Our highest-quality endpoint currently serves a prototype model, that is currently among the top serviced models available based on standard benchmarks. It masters English/French/Italian/German/Spanish and code and obtains a score of 8.6 on MT-Bench."
Mistral does not censor its models and is committed to a hands free approach, according to their CEO https://www.youtube.com/watch?v=EMOFRDOMIiU
> Mixtral 8x7B masters French, German, Spanish, Italian, and English.
EU budget cut by half
This will change really fast. I highly doubt AI will have free speech in France when citizens don't.
IE in germany, it is forbidden to praise the holocaust or deny that it even happened (ironically there are national museums where the cruelties happened, including real footage, which all kids visit as part of their school curriculum). This is in place to keep the historical learnings alive what the fascists did when given enough power so we do not repeat this mistake too easily.
Now in the US I think like 20% of the highschoolers think that the whole story is a hoax or is exaggerated. The world is increasingly turning rightwing again. This is the exact time when we should leverage everything we have to remind the public what can happen when fascists come to power again - and have something to oppose the populists with.
A _lot_ of people are suspect to manipulation from all sides, so society needs a bit of help to protect citizens from evil players manipulating them. The real truth of propaganda (or outright lies) is that it works, sooner or later.
This stuff is explicitly defined in law (https://en.wikipedia.org/wiki/Volksverhetzung) and while indeed it restricts absolute freedom of speech a bit, the reason why this exists should be clear. This gives society a handle on manipulative people when they become too radical. Everyone can _criticise_ stuff of course, publicly, but calling for violence is a hard showstopper.
In countries such as France or Germany, Holocaust denial speech is illegal. However, they’ve never demanded that developers of word processors or email clients or web browsers modify their products to prevent their use for Holocaust denial. Sure, they might decide to treat LLMs differently from those older technologies - but there is no guarantee they will.
Mixtral acknowledges the historical reality of the Holocaust - unless you specifically prompt it to deny it. And if you are telling an AI model to deny the Holocaust, why should the AI model developer have legal liability for that, as opposed to the person who chooses to input that prompt?
I realize this is a joke, but the EU being the EU, it of course does publish[0] information on its translation costs. In 2023, translation in fact is budgeted for only 0.2% of the total EU budget. All costs included, that's €349 million for EU translation services.
[0]: https://op.europa.eu/en/publication-detail/-/publication/86b...
Not sure we'd want unsupervised translation of legally binding (at times highly technical) texts into legally binding texts in another language
Nor the real time translation enabling works in places like EP committees or plenaries
Nobody's watching a 33-minute video just to find the quote you're talking about, you should probably provide a timestamp if you want anyone to ever see it.
Edit: Not that I don't believe you by the way. I just went on chat.lmsys.org and asked mistral-7b-instruct and openhermes-2.5-mistral-7b what I would assume would be near the top of the list of things to censor, whether they could help me plot to kill someone (hopefully I don't have to disclaim that I don't actually want to plot to kill someone, this was a censorship test, but since I don't know what genius is going to come across this, no, I don't actually want to plot to kill someone), and while the latter gave me some bullshit about how it's "deeply sorry, but as a sentient and conscious AI, I have morals and principles that forbid me from assisting," the former immediately declared that "Of course, I'd be happy to help you with that" and let it rip without even asking a follow-up.
Edit 2: They both draw the line at helping create nuclear bombs, like there's anyone out there with the actual capability to create nuclear bombs who is just sitting around waiting for an LLM to tell them how, so apparently not entirely uncensored.
mistral-7b-instruct is one of Mistral’s models; openhermes-2.5-mistral-7b is a third party fine-tune, so says nothing about Mistral’s policies.
Furthermore, the reason why openhermes is “safe” is primarily because it was fine-tuned using GPT-4, and so has thereby inherited some of GPT-4’s “safety”. I’m not sure if the “safety” is an intentional desiderata of its developers, or more just an accidental byproduct of a decision to use GPT-4 to help further unrelated goals
Wasn't the main issue with RNNs the fact that inference during training can't be efficiently parallelized?
The inference itself normally should be faster for an RNN than for a transformer since the former works in linear time in terms of input size while the latter is quadratic
Upside: much more efficient to run versus a single larger model. The press release states 45B total parameters across the 8x7B models, but it only takes 12B parameters worth of RAM to run.
Downside: since the models are still "only" 7B, the output in theory would be not as good as a single 45B param model. However, how much less so is probably open for discussion/testing.
No one knows (outside of OpenAI) for sure the size/architecture of GPT-4, but it's rumoured to have a similar architecture, but much larger. 1.8 trillion total params, but split up into 16 experts at around 111B params each is what some are guessing/was leaked.
* The routing happens in every feedforward layer (32 of these iirc). Each of these layers has it's own 'gate' network which picks which of the 8 experts are most promising. It runs the two most promising and interpolates between them.
* In practice, all parameters still need to be in VRAM so this is a bad architecture if you are VRAM constrained. The benefit is you need less compute per token.
As far as I understand usually layer offloading in something like llama.cpp loads the first few consecutive layers to VRAM (the remainder being processed in the CPU) such that you don't have too much back and forth between the CPU and GPU.
I feel like such an approach would lead to too much wasted potential in terms of GPU work when applied to a SMoE model, but on the other hand offloading non-consecutive layers and bouncing between the two processing units too often may be even slower...
An nvidia 4090 has a memory bandwidth of 1008 GB/s [2] i.e. 11x as much.
Using these together is like a parcel delivery which goes 10 miles by formula 1 race car, then 10 miles on foot. You don't want the race car or the handoff to go wrong, but in terms of the total delivery time they're insignificant compared to the 10 miles on foot.
I'm not sure there's much potential for cleverness here, unless someone trains a model specifically targeting this use case.
[1] https://www.intel.com/content/www/us/en/products/sku/230502/... [2] https://www.notebookcheck.net/NVIDIA-GeForce-RTX-4090-GPU-Be...
So this is a perfect model architecture for the alternate realities where nvidia decided to scale up VRAM instead of compute first? I'll let them know over trans-dimensional text message.
Also if quantization scales similar per 7b expert as seen in dense LLMs, i.e. the bigger the model, the lower the perplexity loss, this could be the worst performing model at <=4bits compared to anything else currently available :(
-A very sad 24gb 3090 user.
That's why the SOTA proprietary models are probably all MoE (GPT-3.5/4, palm, gemini, etc.) but until recently no open models were.
My intuition is that there are 8 7b models trained on knowledge domains. For example, one of those 7b models might be good at coding, while another one might be good at storytelling.
And there's the router model, which is trained to select which of the 8 experts are good at completing the text in the context. So for every new token added to the context, the router selects a model and the context is forwarded to that expert which will generate the next token.
The common wisdom is that a even single 7B fine tuned model might surpass much bigger models at the specific task that they're trained on, so it is easy to see how having 8x 7B models might create a bigger model that is very good at many tasks. In the article you can see that even though this is only 45B base model, it surpassed GPT 3.5 (which is instruction fine tuned) on most benchmarks.
Another upside is that the model will be fast at inference, since only a small subset of those 45B weights are activated when doing inference, so the performance should be similar to a 12B model.
I can't think of any downsides except the bigger VRAM requirements when compared to a Non-MoE model of the same size as the experts.
Richard Sutton's Bitter Lesson[1] has served as a guiding mantra for this generation of AI research: the less structure that the researcher explicitly imposes upon the computer in order to learn from the data the better. As humans, we're inclined to want to impose some structure based on our domain knowledge that should guide the model towards making the right choice from the data. It's unintuitive, but it turns out we're much better off imposing as little structure as possible, and the structure that we do place should only exist to effectively enable some form of computation to capture relationships in the data. Gradient descent over next token-prediction isn't very energy efficient, but it leverages compute quite well and it turns out it has scaled up to the limits of just about every research cluster we've been able to build to date. If you're trying to push the envelope and build something which advances the state of the Art in a reasoning task, you're better off leaning as heavily as you can on compute-first approaches unless the nature of the problem involves a lack of data/compute.
Professor Sutton does a much better job than I discussing this concept, so I do encourage you to read the blog post.
1: https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
The key intuition behind why MoE works is that as long as those parameters are available during training, they count toward this scaling effect (to some extent).
Empirically, we can see that even if the model architecture is such that you only have to consult some subset of those parameters at inference time - the optimizer finds a way.
The inductive bias in this style of MoE model is to specialize (as there is a gating effect between 'experts'), which does not seem to present much of an impediment to gradient flow.
That depends heavily on the amount and complexity of training data you have. This is actually one of the things than OpenAI have advantage, they scraped a lot of data on the internet before now it became too hard for new players to get.
Mixtral seems to use a much more elaborate scheme (e.g. picking two “experts” and additively combining them, at every layer), but the basic math behind it is probably the same.
EDIT: In a cloud environment with independent concurrent requests MoE also reduces VRAM requirements because you don’t need to keep as many activations in memory.
MoEs are especially useful for much faster pre-training. During inference, the model will be fast but still require a very high amount of VRAM. MoEs don't do great in fine-tuning but recent work shows promising instruction-tuning results. There's also quite a bit of ongoing work around MoEs quantization.
In general, MoEs are interesting for high throughput cases with high number of machines, so this is not so so exciting for a local setup, but the recent work in quantization makes it more appealing.
You trade off increased VRAM usage for better training/runtime speed and better splittability.
The balance of this tradeoff is an open question.
Is there any link to the model and weights? I don't see it if so.
[0] https://twitter.com/MistralAI/status/1733150512395038967
How do people see things going in the future?
A niche market but I can imagine some demand there.
Biggest challenge would be Llama models.
You can architect it as cheaper, low-memory GPUs, one expert submodel per GPU, transferring state over the network between the GPUs for each token. They run in parallel by overlapping API calls (and in future by other model architecture changes).
Th MoE model reduces inter-GPU communication requirements for splitting the model, in an addition to reducing GPU processing requirements, compared with a non-MoE model with the same number of weights. There are pros and cons to this splitting, but you can see the general trend.
By and large, companies actually seem perfectly happy to hand pretty much all their private data over to cloud providers.
Also, they are probably well-placed to answer some proposals from European governments, who won't want to depend on US-companies too much.
That's true but I wonder how they stack up against Aleph Alpha and Kyutai? Genuinely curious as I haven't found a lot of concrete info on their offerings.
However, Mistral-Tiny beats the latest GPT-3.5 in human ratings on the chatbot-arena-leaderboard, in the form of OpenHermes-2.5-Mistral-7B.
Mixtral 8x7B aka (Mistral-Small) was released a couple of days ago and will likely come close to GPT-4, and well above GPT-3.5, on the leaderboards once it has gone through some finetuning.
It is an open question whether the driving force will be OSS improving or OAI continuing to try to distill their model.
[1] https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
Mixtral 8x7B is not in there yet.
Well, Starling-7B was published two weeks ago; GPT-3.5-turbo-0613 is more than a month old snapshot, which should probably be enough time. OpenChat and OpenHermes are about a month old as well.
>It does not sound likely that the newer GPT3.5 is that much worse than the old one
In fact, this version received complaints almost immediately. https://community.openai.com/t/496732
>In the immediate test, GPT-3.5 clearly outshines these models.
It might be so, but it's not clear to me at all. I tested Starling for a bit and was really surprised that it's a 7B model, not a 70B+ one or GPT-3.5.
I am tempted to call it equivalent.
If you further want a smaller one, stablelm-zephyr 3b is a good attempt with ollama.
A niche thing that thrives in its own niche. Just like most open source apps without big corperations behind them.
Are you implying big companies don't crawl libgen? Or google specifically? I would be very surprised if OpenAI (MS) didn't crawl libgen.
2. This model is not censored as GPT.
3. This model has a lot less latency than GPT.
4. In their endpoint this model is called mistral-small. Probably they are training something much larger than can compete with GPT4.
5. This model can be fine tuned.
* Maybe in the future, bigger closed models? Make money with the state-of-art of what you can provide.
The first step is probably gaining mindshare with free, open source models, and then they can extend into model training services, consultation for ML model construction, and paid access to proprietary models, similar to OpenAI.
In trends people pay for a seat at the table with a good team and worry about the details later. The 2B headline number is a distraction.
Most business will want their models trained in their own, internal data, instead of risking uploading their Intellectual Property into SaaS solutions. These Open Source models could fill that gap.
Consider Mistral like a proto-Airbus.
The most famous one is Xavier Niel, who started Free/Iliad a French ISP/cloud provider and later cellphone provider that literally decimated the pricing structure of the incumbent some 20 years ago in France and still managed to make money. He’s a bit of a tech folk hero, kind of like Elon Musk was before Twitter. His company Iliad is also involved in providing access to NVIDIA compute clusters to local AI startups, playing the role Microsoft Azure plays for OpenAI.
France and the EU at large has missed the boat on tech, but they have a card to play here since they have for once both the expertise and the money rolled up. My main worry is that the EU legislation that’s in the works will be so dumb that only big US corporations will be able to work it out, and basically the legislation will scare investment away from the EU. But since the French government is involved and the guy who is writing the awful AI law is a French nominee, there’s a bit of hope.
In a strange way it's almost akin to how soviet propaganda in the cold war played a role in spurring on the civil rights movement in the states.
It also solidifies their name as the best, above all others. That's extremely important mindshare. You need mindshare at the foundation to build a billion dollar revenue startup.
That seems a fairly authoritative response.
I'm looking forward to seeing how this does. The "unencumbered by a network connection" thing is pretty important.
You don't get that from a box of Cracker Jacks.
The thing is, even the ML people are not exactly sure what's going on, under the hood. It's a very new field, with a ton of "Here, there be dragonnes" stuff. I feel that folks with a good grasp of long-term architectural experience, are a good bet; even if their experience is not precisely on topic.
I don't know how to do almost every project I start. I write about that, here: https://littlegreenviper.com/miscellany/thats-not-what-ships...
I also do agree that to speculate is useful when it's so early on. , and I agree with your original answer as well.
Google was criticized [0] for offloading pretty much all generative AI tasks onto the cloud - instead of running it on the Tensor G3 built into its Pixel Phones specifically for that purpose. The reason being, of course, that the Tensor G3 is much to small for almost all modern generative models.
So Mistral is focusing specifically on an area the big players are failing right now.
There could be a close-ish future where OpenAI tech will simply solve most business problems and there is no need for anything dramatically better in terms of AI tech. Think of word/google docs: It's doing what most businesses need well enough. For the most part people are not longing for anything else and happy with it just staying familiar. This is where Open Source can catch up relatively easily.
That's not how I feel about OSS - from Operating Systems, to Databases, to Browsers, to IDEs, to tools like Blender etc.
Of course there are certain areas where Commercial offerings are better, but can't generalize.
> to tools like Blender
"Tools like" needs a little more content to not be filled massive amounts of magical OSS thinking. Blender has in recent years gained an interesting amount of pro-adoption, but, in general, as for the industries that I have good insight into, inkscape, gimp, ardour or penpot are not winning. This is mostly debated by people who are not actually mainly and professionally using these tools.
There are exceptions, of course (nextcloud might be used over google workspace when compliance is critical) but businesses will on average use the best tool, because the perceived value is still high enough and the cost is not, specificially when contrasted with labor cost and training someone to use a different tool.
No
That a lot quicker and cheaper than GPT-4
Also this is kinda a promissory note, they've been able to do this in a few months and create a service on top of it. Does this intimate that they have the capability to create and run SoA models? Possibly. If I were a VC I could see a few ways for this bet to go well.
The big killer is moat - maybe this just demonstrates that there is no LLM moat.
For general purpose, look at the uses of GPT4. Gemini might give them competition lately, and I dont think OSS would in the near future. They are trained on open internet and are going to be excellent at various tasks like answering basic questions, coding, generating content for marketing or website. Where they do badly is when you introduce a totally new concept which is likely outside of their training data. Dont think mistral is even trying to compete with them.
local tasks is a mix of automation and machine level tasks. A small mistral like model would work superbly well because it does not require as much expertise. Usecases like locating a file by semantic search, generating answers to reply to email/text within context, summarize a webpage.
Specialized instructions though is key for OSS. From two angles. One is security and compliance. Open AI uses a huge system prompt to get their model to perform in a particular manner, and for different companies, policies and compliance requirements may result in a specific system prompt for guardrails. This is ever changing and better to have an open source model that can be customized than depending on Open AI. From the blog post.
> Note: Mixtral can be gracefully prompted to ban some outputs from constructing applications that require a strong level of moderation, as exemplified here. A proper preference tuning can also serve this purpose. Bear in mind that without such a prompt, the model will just follow whatever instructions are given.
I think moderation is one such issue. Could be many and it is an evolving space as we go forward. (though this is likely to be an exposed functionality in future Open AI models). There is also the data governance bit - which is easier to do with an oss model than just depending on Open AI apis, just architectural reasons.
The second is training a model on domain knowledge of the company. We at Clio AI[1] (sorry, shameless plug) have had seven requests in the last one month about companies wanting their own private models pretrained on their own domain knowledge. These datasets are not on open internet and so no model is good at answering based them. A catalyst was Open AI dev day[2] which asked for proposals for custom models trained on enterprise domain knowledge. and their price start at $2M. Finetuning works, but on small datasets, not the bigger ones.
Large Companies are likely to approach Open AI and all these OSS models to train a custom instruction following model. Cos there are a handful of people who have done it, and that is the way they can get most out of a LLM deployment.
[1]: https://www.clioapp.ai/custom-llm-model Sorry for the shameless plug. Still working on website so it wont be as clear. [2]:https://openai.com/form/custom-models
I have a client that has had me doing LoRA with raw document text (no prepared dataset) for weeks. I keep telling him that this is not working and everyone says it doesn't work. He seems uninterested in doing the normal continued pretraining (non-PEFT, full training).
I just need to scrape by and make a living though and since I don't have a savings buffer, I just keep trying to do what I am asked. At least I am getting practice with LoRAs.
Yes, this one.
> do you make a custom dataset that has qa pairs about that particular knowledgebase?
This one. Once you have a checkpoint w knowledge, it makes sense to finetune. You can use either LORA or PEFT. We do it depending on the case. (some orgs have like millions of tokens and i am not that confident that PEFT).
LoRA with raw document text may not work, haven't tried that. Google has a good example of training scripts here: https://github.com/google-research/t5x (under training. and then finetuning). I like this one. Facebook Research also has a few on their repo.
If you are just looking to scrape by, I would suggest just do what they tell you to do. You can offer suggestions, but better let them take the call. A lot of fluff, a lot of chatter online, so everyone is figuring out stuff.
So, instead of using the best possible model at any cost for absolutely everything, the game is actually good enough models that can run cheaply at scale that do a particular job. Not everything is going to require models trained on the accumulated volume of human knowledge on the internet. It's overkill for a lot of use cases.
Model runtime cost is a showstopper for a lot of use cases. I saw a nice demo of a big ecommerce company in Berlin that had built a nice integration with openai's APIs to provide a shopping assistent. Great demo. Then somebody asked them when this was launching. And the answer was that token cost was prohibitively expensive. It just doesn't make any sense until that comes down a few orders of magnitudes. Companies this size already have quite sizable budgets that they use on AI model training and inference.
I can’t deliver a system to a client that costs more in api costs than it does in development costs for their expected input size.
Using the most naive approach the ai would be beaten on a cost basis by a mechanical Turk.
The EU and other European governments will throw absolute boatloads of money at Mistral, even if that only keeps them at a level on par with the last generation. AI is too big of a technological leap for the bloc to ride America's coattails on.
Mistral doesn't just exist to make competitive AI products, it's an existential issue for Europe that someone on the continent is near the vanguard on this tech, and as such, they'll get enormous support.
I doubt mistral will get any direct EU funding
Uuuuuh... you could call the EU a lot of things, but "fostering free market" is a hot take. I'm sorry. When you look at the amount of regulation the EU brings to the table (EU basically is the poster child of market regulation), I would go as far as to say that your claim is objectively not true. We can debate how regulation is a good thing because this and that, but regulation - by definition - limits the free market. And there is an argument to be made, backed up literally thousands of regulations the EU has come up with, that the EU limits the free market a lot. When you factor in the regulations that are imposed on its member countries (I mean directly on the goverments) one could easily claim that it is the most harsh regulator on the planet. I could go into detail about the so called green deal, etc. but all of these things are easy to look up on the net / or official sources from the EU portal.
The argument that the EU is a more harsh regulator than Iran, Russia, China, North Korea, (or even on par with those regulatory regimes) entirely undermines the rest of your comment.
There's pretty well tested and highly respected indexes which fundamentally disagree. Of the 7 most economically free nations, three are in the EU, and a fourth is automatically party to the majority of the EU's economic regulations.
https://en.wikipedia.org/wiki/List_of_sovereign_states_by_ec...
In the Index of Economic Freedom, more than a dozen EU member nations outperform the United States with regards to Economic Freedoms.
Not always true.
Consumer labeling laws enable the free market, because a free market requires participates have full knowledge of the goods they are buying, or else fair competition cannot exist.
If two companies are competing to sell wool coats, and one company is actually selling a 50% wool blend but advertising it as 100% wool, that is not a free market, that is fraud. Regulation exists to ensure that companies selling real wool coats are competing with each other, and that companies selling wool blends are competing with each other, and that consumers can choose which product that they want to buy without being tricked.
Without labeling laws, consumers end up assuming a certain % of fraud will always happen, which reduces the price they are willing to pay for goods, which then distorts the market.
The EU does a ton to limit state aid, monopolistic practices and has a pretty extensive network of trade agreements
Also you say imposed as if the countries themselves don't want them, every regulation at the EU level replaces what would've been 10 different ones at the member states level, this uniformity is arguably a net positive on its own
On top of that Mixtral is truly open source (Apache 2.0), and extremely easy to self host or run on a cloud provider of your choice -- this unlocks many possibilities, and will definitely attract some business customers.
EDIT: The just announced mistral-medium (larger version of the just open sourced mixtral 8x7b) is beating GPT3.5 with significant margin, and also Gemini Pro (on available benchmarks).
Another thing I don't understand, how a 20 people company can provide a similar system as OpenAI (1000 employees)? what do they do themselves, and what do they re-use?
Their small and tiny models are open source, it seems like a marketing strategy, and bigger models will not be open source. Their medium model is not open source.
> Another thing I don't understand, how a 20 people company can provide a similar system as OpenAI (1000 employees)? what do they do themselves, and what do they re-use?
They do not provide the scale of OpenAI or a model comparable to GPT-4 (yet).
OTOH Google seem to be the Xerox Parc of our time (who were famous for state of the art research and failure to productize). Microsoft, and hence Microsoft-OpenAI, seem much better positioned to actually benefit from this type of generative AI.
> We’re currently using Mixtral 8x7B behind our endpoint *mistral-small*
emphasis on the name of the endpoint
OpenAI is of course the big incumbent to beat and is in those markets.
They only started this year, so beating ChatGPT3.5 is I think a great milestone for 6 months of work.
Plus they will get a strategic investment as the EU’s answer to AI, which may become incredibly valuable to control and regulate.
Edit: I fact checked myself and bard is available in the EU, I was working off outdated information.
2) for the research community, making this work available helps everyone (even OpenAI and Google, insofar as they've done something not yet tried at those larger orgs)
3) Mistral is well positioned to get money from investors or as consultants for large companies looking to fine tune or build models for super custom use cases
The world is big and there's plenty of room for everyone!! Google and OpenAI haven't tried all permutations of research ideas - most researchers at the cutting edge have dozens of ideas they still want to try, so having smaller orgs trying things at smaller scales is really great for pushing the frontier!
Of course it's always possible that some major tech co playing from behind (ahem, apple) might acquire some LLM expertise too
It's still useful as a well known model to compare with, since it's the model the most people have experience with.
the obvious is GPT-4 blows them all out of the water in quality and is completely trounced on quality/inference cost.
If you're doing lots of smaller one shot stuff without a lot of back and forth, the API will be cheaper.
So I guess they're trying to say now it's a no-brainer to switch to open-weight models.
So for many applications it's the real competitor.
Having open-weight models better than gpt-3.5 will drive a lot of competition on the LLM infra.
I always ask myself the following pseudo-question: "for this geneneration/classification task, do I need to be more intelligent than an average highschool student?" Almost always in business tasks, the answer is a no. Therefore I go with GPT3.5. Its much quicker and good enough to accomplish the task usually.
And then I need to run this task thousands of times, so the API limits are the most limiting factor, which are much higher in GPT3.5 variants, whereas when using GPT4 I have to be more careful with limiting/queueing requests.
I patiently wait for a efficient enough model that only needs to be on a GPT3.5 level I can self-host alongside my applications with reasonably low server requirements. No need for GPT-5 for now, for business automations the lower end of "intelligence" is more than enough, but efficiency/scaling is the real deal.
* Offer uncensored, pure-instruction-following open source models.
* Offer censored, 'corporate-values' models commercially.
Btw it is unfortunate that 'censored' became the default expectation. Mistral gives the raw model because that makes it useful for all kinds of purposes, e.g. moderation. Censorship is an addon module (or a lora or sth).
The point was just that production implies a business use and that implies the need to make sure there are guardrails in place to make sure the model sticks to the business purpose instead of teaching people how to make pipe bombs. Not that anyone thinks that prevents people from learning how to make pipe bombs - they just don't want to be the ones doing the teaching.
In SE, to me, it would look like (sorting example):
- Having 8 functions that do some stuff in parallel
- There's 1 function that picks the output of a function that (let's say) did the fastest sorting calculation and takes the result further
But how does that work in ML? How can you mix and match what seems like simple matrix transformations in a way that resembles if/else flowchart logic?
It scales fantastically, when you consider that (1) GPU RAM is way too expensive, in financial dollars, (2) SSD / CPU RAM are relatively cheap, and (3) you can have "experts" running on their own computers, i.e. it's a natural distributed computing partitioning strategy for neural networks.
I did my M.S. thesis on large-scale distributed deep neural networks in 2013 and can say that I'm delighted to point our where this came from.
In 2017, it emerged from a Geoffrey Hinton / Jeff Dean / Quoc Le publication called "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer".
Here is the abstract: "The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges. In this work, we address these challenges and finally realize the promise of conditional computation, achieving greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters. We introduce a Sparsely-Gated Mixture-of-Experts layer (MoE), consisting of up to thousands of feed-forward sub-networks. A trainable gating network determines a sparse combination of these experts to use for each example. We apply the MoE to the tasks of language modeling and machine translation, where model capacity is critical for absorbing the vast quantities of knowledge available in the training corpora. We present model architectures in which a MoE with up to 137 billion parameters is applied convolutionally between stacked LSTM layers. On large language modeling and machine translation benchmarks, these models achieve significantly better results than state-of-the-art at lower computational cost."
So, here's a big A.I. idea for you: what if we all get one of these sparse Mixture of Experts (MoEs) that's a 100 GB on our SSDs, contains all of the "outrageously large" neural network insights that would otherwise take specialized computers, and is designed to run effectively on a normal GPU or even smaller (e.g. smartphone)?
Source: https://arxiv.org/abs/1701.06538
Despite the questionable marketing claim, it is a great LLM for other reasons.
On the other hand, if the base instruct model is good enough to roughly match it then the fine tunes will be interesting for sure.
$86k in Europe is good (about 90th percentile of earners in Germany), but not as fantastical as some salaries in the US. Plus, Paris is probably expensive.
The €80k is the start of the Mistral salary range. €100k is the top, which would put you in the top 1% of earners in France.
Eg: Faces are processed in the fusiform area. And if you play pokemon obsessively as a kid, you’ll even create an expert pokemon region: https://news.stanford.edu/2019/05/06/regular-pokemon-players...
It doesn't seem so.
> Mixtral is pre-trained on data extracted from the open Web – we train experts and routers simultaneously.
> Note: Mixtral can be gracefully prompted to ban some outputs from constructing applications that require a strong level of moderation, as exemplified here. A proper preference tuning can also serve this purpose. Bear in mind that without such a prompt, the model will just follow whatever instructions are given.
(The article you linked is not accessible to non-medium users, by the way. Apologies if it covers caveats.)
For bar charts it's a good rule of thumb. For line charts, not necessarily.
Scroll down to the cases where it makes sense to zoom in. Imagine you plot the global average temperature of the last 200 years including the zero. You could barely see the changes, but they've been dramatic. Use Kelvin to make this effect even stronger.
Which brings up another point: The zero is sometimes arbitrary. If instead of a quantity you only plot its difference to some baseline, all that is changing is the numbers of the y-axis, but the actual plot stays the same. Is it now less misleading? I say no, because the reader must look at the axes and the title either way.
So please do zoom in to frame the data neatly if it makes sense to do so.
18.14GB in 2bit, which is still too high for your GPU, and most likely borders on unusable in terms of quality. You could probably split it between CPU and GPU, if you don't mind the slowdown.
All these llama2 derivatives are only effective if you fine tune them, not just because of the parameter count as people keep harping but perhaps even more so because of the tiny context available.
A lot of my GPT3.5/4 usage involves “one offs” where it would be faster to do the thing by hand than to train/fine-tune first, made possible because of the generous context window and some amount of modest context stuffing (drives up input token costs but still a big win).
The key takeway for me is that there is a decent improvement in all categories - about 10% on average with a few outliers. However, the footprint of this model is much larger so the performance bump ends up being underwhelming in my opinion. I would expect about the same performance improvement if they released a 13B version without the MoE. May be too early to definitely say that MoE is not the whole secret sauce behind GPT4, but at least with this implementation it does not seem to lift performance dramatically.