Mistral: Our first AI endpoints are available in early access
mistral.ai
mistral.ai
This is a tiny company (appears 30 or so people?) that just scored a 2B valuation, produced easily the most performant 7B model and a 7B*8 MOE model that performs at the level of a 70B requiring the inference power of a 14B.
I feel this could be a potential bigger threat to OpenAI than Google or Anthropic. I gather with the huge recent investment they'll be able to a) scale out to a reasonable traffic load in the near future and b) attract the best and brightest researchers put off w/ various chest puffing and drama that has been front and center in this industry.
They managed to put out the gold standard for 7B models in like 6 months, and are quickly moving up the scale.
I mocked the funding round back in March as being signs of hype, ($300m for a team of 3 with just an idea?) but clearly I didn't know the details. Really remarkable execution.
They may well be on their way to eat all use cases that don't need gpt-4 performance and hopefully soon tackle the big leagues as well. Exciting times!
LLMs are neat, but likely wont have a place in an AGI.
Maybe in 1970 someone was arguing that Volvo is handicapped because Swedish safety regulations have become so stringent that their cars can’t compete on price elsewhere. Turns out it was exactly Volvo’s strength to be ahead of the curve on regulation.
I've seen a few graphs showing staffing counts and budgets for US regulatory agencies have grown exponentially since the 1970s (adjusted for inflation), while it was relatively flat before then in the 60s. Economic regulation growth in the US has grown steadily while social regulations (workplace, climate, healthcare, transportation etc) grew dramatically. Adding in TSA / Homeland security and it's a hockey stick.
https://regulatorystudies.columbian.gwu.edu/sites/g/files/za...
I'd be curious to see similar charts for the EU.
[0]https://www.ft.com/content/9339d104-7b0c-42b8-9316-72226dd4e...
That's a weird analogy. I don't remember Volvo ever being that big. If anything they're bigger now than they used to be (I mean: I don't remember a world with lots of Volvo when I was a kid... I may be wrong though). And with about 650 000 vehicles sold per year they're not in the Top 25 car manufacturer worldwide by number of cars sold.
Also, I'm not sure how much the EU is really going to matter for the future of AI.
You don't really need to conform to their regressive tendencies if you don't want to do business with the EU. Or you could dumb your stuff down for them if it ever comes down to it.
Generally, it just wouldn't be worth it to care too much about what the EU thinks.
It's kind of hard to tell what it is from both the blog post and their homepage. So only people really familiar with AI will grasp the relevance. But your comment certainly helps.
Mixtral of experts - https://news.ycombinator.com/item?id=38598559 - Dec 2023 (272 comments)
Mistral-8x7B-Chat - https://news.ycombinator.com/item?id=38594578 - Dec 2023 (69 comments)
Mistral AI Valued at $2B - https://news.ycombinator.com/item?id=38593616 - Dec 2023 (221 comments)
Mistral's mixtral-8x7B-32kseqlen on Vercel - https://news.ycombinator.com/item?id=38584179 - Dec 2023 (30 comments)
French AI startup Mistral secures €2B valuation - https://news.ycombinator.com/item?id=38580758 - Dec 2023 (76 comments)
Mistral "Mixtral" 8x7B 32k model [magnet] - https://news.ycombinator.com/item?id=38570537 - Dec 2023 (236 comments)
I know these are not all the same exact story but the discussions end up being more or less generically the same, so we can call them all (or most of them) quasidupes.
My argument is opening up the commercial offering is perhaps the biggest story yet simply because of all excitement this company has been able to generate.
[1] I would expect real world-performance gap to be even larger if Mistral 7B is anything to go by. The fact that safety filters are opt-in is a huge benefit (even for safe applications).
Pretty sad for Google if their next big AI thing is already being beaten by a small company with a tiny fraction of their resources.
As for AI they have a lot of world-class talent but they somehow cannot transform that into a coherent product even though they did amazing stuff on the research side. It has been brought up many times on HN on how Google has a product-design problem and it seems to be true.
Google can't do anything about it either - consumers will prefer the chat approach and will go where it is offered. Google has to offer it, or people will go to ChatGpt and Bing, but at the same time they lose a lot of money each time people opt for the chat summary.
It is way expensive for Bing too, but they don't care because their market is growing and since the only thing people used to search Bing for was how to download Chrome and how to change the default browser, they aren't losing a lot this way.
I think search products will be OK - it will just be stranger than we expect.
Bing going legit and genuinely threatening google wasn’t on my bingo card for sure.
But those ads would be compelling and more valuable. There is no reason to believe the current inference LLM architecture is the optimal one (as shown by Mistral's MoE) - future inference will likely be cheaper than it is currently.
The sentiment that Google's goose is cooked I'm seeing in this thread seems like wish-fulfillment to me. Especially considering that the upstarts also don't have a clear road to profitability
That’s a large and difficult problem obviously and one Google is clearly well-positioned for if they can ever catch up on the model front.
Perhaps even sooner, given the pace tech is evolving nowadays comparing to when Google took over Yahoo's place.
Re: "distracting other companies", there's however this famous essay by Joel Spolsky (2002):
https://www.joelonsoftware.com/2002/01/06/fire-and-motion/
> Think of the history of data access strategies to come out of Microsoft. ODBC, RDO, DAO, ADO, OLEDB, now ADO.NET – All New! Are these technological imperatives? The result of an incompetent design group that needs to reinvent data access every goddamn year? (That’s probably it, actually.) But the end result is just cover fire. The competition has no choice but to spend all their time porting and keeping up, time that they can’t spend writing new features. Look closely at the software landscape. The companies that do well are the ones who rely least on big companies and don’t have to spend all their cycles catching up and reimplementing and fixing bugs that crop up only on Windows XP. The companies who stumble are the ones who spend too much time reading tea leaves to figure out the future direction of Microsoft. People get worried about .NET and decide to rewrite their whole architecture for .NET because they think they have to. Microsoft is shooting at you, and it’s just cover fire so that they can move forward and you can’t, because this is how the game is played, Bubby.
In fact Google did a lot against it. Google search was clean, fast, and with no ads, GMail had one of the best spam filters, and Google algorithms were highly resistant to SEO of the time. Time have passed, Google is now part of the problem, but it is not alone there, and for some time, Google really did good.
It's not really "available" though, is it? I won't buy any PR benchmarks until the models are publicly available, there's too much variation based on how the models need to toned back for safety reasons before being released to the public
Pricing has been released too.
Per 1 million output tokens:
Mistral-medium $8
Mistral-small $1.94
gpt-3.5-turbo-1106 $2
gpt-4-1106-preview $30
gpt-4 $60
gpt-4-32k $120
This suggests that they’re reasonably confident that the mistral-medium model is substantially better than gpt3-5
Mistral-small seems to be the most direct competitor to gpt-3.5 and it’s cheaper (1.2 eur / million tokens)
Note: I’m assuming equal weight for input and output tokens, and cannot see the prices in USD :/
(That one is a curious case. I once spent some time trying to figure out why no major photo app seems to support manually tagging faces, which is a mind-dumbingly obvious feature to support, and which was something supported by software a decade or so ago. I couldn't find anything definitive; there's this eerie conspiracy of silence on the topic, that made me doubt my own sanity at times. Eventually, I dug up hints that some EU ruling/regs related to facial recognition led everyone to remove or geolock this feature. Still nothing specific, though.)
Inference on a quantized model is faster, not slower.
However, I have no idea how practical it is to run a LLM on a phone. I think it would run hot and waste the battery.
Though I guess that's not for lack of trying.
Mistral may just be aiming for a more sustainable price for the long run.
How did you reach the conclusion? Maybe they are counting on people paying extra just to prevent vendor lockdown.
I also know that because it’s open source, if I ever have a need to, I can host it on my own servers. Currently I don’t have that need, but it’s nice to know that it’s in the cards.
OpenAI just went through an existential crisis where the company almost collapsed. They are also quite unreliable. For some use cases, I'll take a service that does slightly worse on outputs, but much better on reliability. For example, if I'm building a customer service chat bot, it's a pretty big deal if the LLM backend goes down. With an open-source model, I can build it using the cloud provider. If they are a reliable host, i'll probably stick with them as i grow. If not, I always have the option of running the model myself. This alleviates a lot of the risk.
If you carefully craft and evaluate your more complex prompts against a closed model... and then that model is retired, you need to redo that process.
A lot of people were burned when OpenAI withdrew Codex, for example. I think that was a poor decision by OpenAI as it illustrated this exact risk.
If the hosted model you are using is open, you have options for continuing to use it should the host decide to stop offering it.
I just did some napkin math, looks like inference on a 30B model with a GTX 4090 should get you about 30 tokens/sec [1], or 100k tokens/hour.
Considering such systems consume about 1 kW, that's about 10 kWh/1M tokens.
Based on the current cost of electricity, I don't think anyone could get below 2 ~ 4 $ per 1M token for a 30B model.
[1] https://old.reddit.com/r/LocalLLaMA/comments/13j5cxf/how_man...
So, 10kwh could be a lot less than what you cite. That's also how grid operators make money. They generate cheaply and sell with a nice margin. Prices are determined by the most expensive energy sources on the grid in some markets (coal, nuclear, etc.). So, that pricing doesn't reflect actual cost for renewables, which is typically a lot lower than that. Anyone consuming large amounts of energy will be looking to cut their cost. For data centers that typically means investing in energy generation, storage, and efficient hardware and cooling.
That might be a great opportunity for cheap LLMs too.
If you consider 600w for the entire system, that's only 6kWh/1M token, for me 6kWh @0.2USD/kWh is 1.2USD/1M tokens.
And that's without the power efficiency improvements that an H100 has over the 4090. So I think 2$/1M should be achievable once you combine the efficiencies of H100s+batching, etc. Since LLM's generally dwarf the network delay anyway, you could host in places like washington for dirt cheap prices (their residential prices are almost half of what I used for calculations)
Here's how it works in reality:
https://docs.mystic.ai/docs/mistral-ai-7b-vllm-fast-inferenc...
https://www-files.anthropic.com/production/images/model_pric...
When I try to access: “Access to our API is currently invitation-only, but we'll let you know when you can subscribe to get access to our best models.”
Is there any information if this embedding model is or will be open source?
I like it, also made me laugh
from: https://twitter.com/yupiop12/status/1734137238177698106
> 2023-10-21: CUDA support in the Windows version, mistral model support. Speculative sampling is supported. BNF grammar and JSON schema sampling.
> mistral_7B_instruct_q4 - 3.9GB - Mistral 7B chat model
This is interesting. This model outperforms ChatGPT 3.5. I'm not sure what type of model it is, and it is not open-sourced.
> Mistral-tiny. Our most cost-effective endpoint currently serves Mistral 7B Instruct v0.2, a new minor release of Mistral 7B Instruct. Mistral-tiny only works in English. It obtains 7.6 on MT-Bench. The instructed model can be downloaded here.
"download here" link is to v0.1 [0]. Oversight or are they holding back the state of the art tiny model?
Does anyone know the kind of actual infrastructure something like gpt4-32k actually run on?
I mean when I actually type something in the prompt, what actually happens behind the scenes?
Is the answer computed on a single NVidia GPU?
Or is it dedicated H/W not known to the general public?
How big is that GPU?
How much RAM does it have?
Is my conversation run by a single GPU instance that is dedicated to me or is that GPU shared by multiple users?
If the latter, how many queries per seconds can a single GPU handle?
Where is that GPU?
Does it run in an Azure data center?
Is the API usage cost actually reflective of the HW cost or is it heavily subsidized?
Is a single GPU RAM size the bottleneck for how large a model can be?
Is any of that info public ?
Also we can probably assume the pricing is likely to be somewhat in proportion to the cost to run (possibly subsidised to gain market, but they are unlikely to be taking a giant/unsustainable loss per query here, particularly as they seem to announce price decreases when they increase model performance).
Most likely given that one of their open positions for a GPU programmer includes
> high technical competence for writing custom CUDA kernels and pushing GPUs to their limits.
Edit: only narrows it down to NVidia hardware, IDK if single GPU or not.
FWIW, I most enjoyed the 29TB machine demo at the end.
All these llama2 derivatives are only effective if you fine tune them, not just because of the parameter count as people keep harping but perhaps even more so because of the tiny context available.
A lot of my GPT3.5/4 usage involves “one offs” where it would be faster to do the thing by hand than to train/fine-tune first, made possible because of the generous context window and some amount of modest context stuffing (drives up input token costs but still a big win).
What are you basing this observation on; personal experience, or is there a benchmark somewhere confirming it?
I tried using it for "business document" use cases but have ran into this with code as well; the latter might be a better explanation given where we're having this discussion. If you only need the llm to retain the general shape of your inputs so it can reuse them to influence the output, the sliding context is fine. But if you need it to actually reuse code verbatim from the input that you fed it (or to remember the api calls and their surrounding context verbatim to recall from a sample of just one that this api must be called before that api, when the prompt includes instructions to that effect) the "decomposition" of the input tokens with the sliding model is insufficient and the llm completely fails at the assigned task.
It refers to the original Mistral 7B though not the new Mixtral fwiw
Again, sufficient training can overcome these limitations. But that's only for cases where the corpus of input documents is static or at least contains significant reuse.
The thing that makes me a bit sad is how announcements show the benchmarks but they way they test is tweaked to make their metrics favorable. They aren't apple to apple benchmarks across different paper publications.
Super grateful that they openly share the weights and code with Apache license.
25 shot - is having 25 tries and selecting the best answer.
Is anyone working on an open benchmark where they take the major models and compare them apples to apples.
Can it run on 1 GPU and swap between experts.
Then they'd be differing opinions?