Mistral AI Valued at $2B
unite.ai
unite.ai
Another company worthy of some hype is 01.AI which released their Yi-34B model. I have been running Yi locally on my Mac (use “ ollama run yi:34b”) and it is amazing.
Hype away Mistral and 01.AI, hype away…
You can play with them, tune them, and download the weights
It isn’t exactly the same as open source because weights != source code, but it is close in the sense that it is editable
IMO we just don’t have great tools for editing LLMs like we do for code, but they are getting better
Prompt engineering, RAG, and finetuning/tuning are effective for editing LLMs. They are getting easier and better tooling is starting to emerge
I noticed that gpt3.5 is practically useless to me (either wrong or too generic), while gpt4 provides a decent answer 80% of the time.
Of course, GPT-5 is expected soon, so there's a moving target. And I can't see myself using GPT-4 much after GPT-5 is available, if it represents a significant improvement. We are quite far from "good enough".
It can make our jobs a lot easier or it can take our jobs.
We are trying to keep SWE salaries up, and lowering the barrier to entry will drop them.
Maybe a more advanced type of model they'll invent in the next years. Who knows... But GPT-like models? Nah, they won't write useful code applicable in prod without supervision by an experient engineer.
Step 1: get a billion dollars.
That’s your main trade secret.
Humans learn a lot of things from very little input. Seems to me there's no reason, in principle, that AIs could not do the same. We just haven't figured out how to build them yet.
What we have right now, with LLMs, is a very crude brute-force method. That suggests to me that we really don't understand how cognition works, and much of this brute computation is actually unnecessary.
It's precisely because we don't know how to build these LLMs cheaply that one must so spend so much money to build them.
Transistors used to cost a billion times more than they do now [1]. Do you have any reason to suspect AIs to be different?
[1] https://spectrum.ieee.org/how-much-did-early-transistors-cos...
However you would still need billions of dollars if you want state of the art chips today, say 3nm.
Similarly, LLM may at some point not require a billion dollars, you may be able to get one, on par or surpass GPT4, easily for cheap. The state of the art AI will still require substantial investment.
And also takes 8 hours of sleep per day, and are mostly worthless for the first 18 years. Oh, also they may tell you to fuck off while they go on a 3000 mile nature walk for 2 years because they like the idea of free love better.
Knowing how birds fly ready doesn't make a useful aircraft that can carry 50 tons of supplies, or one that can go over the speed of sound.
This is the power of machines and bacteria. Throwing massive numbers at the problem. Being able to solve problems of cognition by throwing 1GW of power at it will absolutely solve the problem of how our brain does it with 20 watts in a faster period of time.
The original point was that an “AI” might become so advanced that it would be able to describe how to create a brain on a chip. This is flawed for two main reasons.
1. The models we have today aren’t able to do this. We are able to model existing patterns fairly well but making new discoveries is still out of reach.
2. Any company capable of creating a model which had singularity-like properties would discover them first, simply by virtue of the fact that they have first access. Then they would use their superior resources to write the algorithm and train the next-gen model before you even procured your first H100.
According to [1] a 70B model needs $1.7 million of GPU time.
And when you spend that - you don't know if your model will be a damp squib like Bard's original release. Or if you've scraped the wrong stuff from the internet, and you'll get shitty results because you didn't train on a million pirated ebooks. Or if your competitors have a multimodal model, and you really ought to be training on images too.
So you'd want to be ready to spend $1.7 million more than once.
You'll also probably want $$$$ to pay a bunch of humans to choose between responses for human feedback to fine-tune the results. And you can't use the cheapest workers for that, if you need great english language skills and want them to evaluate long responses.
And if you become successful, maybe you'll also want $$$$ for lawyers after you trained on all those pirated ebooks.
And of course you'll need employees - the kind of employees who are very much in demand right now.
You might not need billions, but $10M would be a shoestring budget.
[1] https://twitter.com/moinnadeem/status/1681371166999707648
This just screams to me that we don’t have a clue what we’re doing. We know how to build various model architectures and train them, but if we can’t even roughly predict how they’ll perform then that really says a lot about our lack of understanding.
Most of the people replying to my original comment seem to have dropped the “in principle” qualifier when interpreting my remarks. That’s quite frustrating because it changes the whole meaning of my comment. I think the answer is that there isn’t anything in principle stopping us from cheaply training powerful AIs. We just don’t know how to do it at this point.
You can't replace those types of LLM with a human, the same way you can't replace Google Search (or GitHub Search) with a human.
Acquiring and preparing that data may end up being the most expensive part.
But creating a base model is out of reach. You need an order of probably hundreds of millions of $$ (if not billion) to get close to GPT 4.
In terms of "secret sauce" it's 95% data quality and 5% architectural choices.
so the ordering is probably data, HW, LLM model
This also fits the general ordering of
data = all human knowledge HW = integrated complexity of most technologists LLM = small team
Still requires the small team to figure out what to do with the first two, but it only happened now because the HW is good enough.
LLMs would have been invented by Turing and Shannon et al. almost certainly nearly 100 years ago if they had access to the first two.
Model merging can create truly unique models. Love to see shit from ghost in the shell turn into real life
Yes training a new model from scratch is expensive, but creating a new model that can’t be replicated by fine tuning is easy
There is indeed already open source models rivaling ChatGPT-3.5 but GPT-4 is an order of magnitude better.
The sentiment that GPT-4 is going to be surpassed by open source models soon is something I only notice on HN. Makes me suspect people here haven't really tried the actual GPT-4 but instead the various scammy services like Bing that claim they are using GPT-4 under the hood when they are clearly not.
HNs funny right now because LLMs are all over the front page constantly, but there's a lot of HN "I am an expert because I read comments sections" type behavior. So many not even wrong comments that start from "I know LLaMa is local and C++ is a programming language and I know LLaMa.cpp is on GitHub and software improves and I've heard of Mistral."
And this is most noticiable if you ask anything that is not in English-American-ish.
I'm in agreement with you, I've been following this field for a decade now and GPT-4 did seem to cross a magical threshold for me where it was finally good enough to not just be a curiosity but a real tool. I try to test every new model I can get my hands on and it remains the only one to cross that admittedly subjective threshold for me.
The early information I see implies it is above. Mind you, that is mostly because GPT-3 was comparatively low: for instance its 5-shot MMLU score was 43.9%, while Llama2 70B 5-shot was 68.9%[0]. Early benchmarks[1] give Mixtral scores above Llama2 70B on MMLU (and other benchmarks), thus transitively, it seems likely to be above GPT-3.
Of course, GPT-3.5 has a 5-shot score of 70, and it is unclear yet whether Mixtral is above or below, and clearly it is below GPT-4’s 86.5. The dust needs to settle, and the official inference code needs to be released, before there is certainty on its exact strength.
(It is also a base model, not a chat finetune; I see a lot of people saying it is worse, simply because they interact with it as if it was a chatbot.)
[0]: https://paperswithcode.com/sota/multi-task-language-understa...
[1]: https://github.com/open-compass/MixtralKit#comparison-with-o...
It's not there yet, but its waaaay closer than the plain Mistral chat release.
It can be useful, but I can see how it'll generate a class of lazy coders who can't think by themselves and just try to get the answer from ChatGPT. An amplified Stack Overflow syndrome.
How does it compare to other models? and with chatgpt in particular?
But you need to run the top Yi finetunes instead of the vanilla chat model. They are far better. I would recommend Xaboros/Cybertron, or my own merge of several models on huggingface if you want the long context Yi.
It's better for the AI ecosystem as a whole to incentive AI startups to make a business through good and open software instead of building moats and lock-in ecosystems.
It’s more akin to a SaaS company releasing a compiled binary that usually runs on their server. Better than nothing, but not exactly in the spirit of open source.
This doesn’t seem like a pedantic distinction, but I suppose it’s up to the community to agree or disagree.
A compiled binary is a bad metaphor because it gives the implication that Mistral-7B is an as-is WYSIWIG project that's not easily modifiable. In contrast, there have been a bunch of new powerful new models created by modifying or finetuning Mistral-7B such as Zephyr-7B: https://huggingface.co/HuggingFaceH4/zephyr-7b-beta
The better analogy to Mistral-7B is something like modding Minecraft or Skyrim: although those games are closed source themselves, it has enabled innovations which helps the open-source community directly.
It would be nice to have fully open-source methodologies but lacking them isn't an inherent disqualifier.
If you want to reproduce the training pipeline, you couldn't do that even if you wanted to because you don't have access to thousands of A100s.
"Source code is defined as the preferred form of the program for making changes in. Thus, whatever form a developer changes to develop the program is the source code of that developer's version."
According to the Open Source Definition:
"The source code must be the preferred form in which a programmer would modify the program. Deliberately obfuscated source code is not allowed. Intermediate forms such as the output of a preprocessor or translator are not allowed."
LLM models are usually modified by changing the model weights directly, instead of retraining the model from scratch. LLM weights are poorly understood, but this is an unavoidable side effect of the development methodology, not deliberate obfuscation. "Intermediate" implies a form must undergo further processing before it can be used, but LLM weights are typically used directly. LLMs did not exist when these definitions were written, so they aren't a perfect fit for the terminology used, but there's a reasonable argument to be made that LLM weights can qualify as "source code".
They're understood based on knowing the training process though, and a developer working on them would want to have the option of doing a partial or full retraining where warranted.
Has Nvidia valued the company at 1B? Say their margin is 80% on the sales. So Nvidia has lost some cashflow and $20M for that 10%. Has Nvidia valued the company at $200M?
1) Is all the money spent up front? Or does it trickle back in over a few years? Cash flow might be impacted more than implied, but I doubt this is much of an issue.
2) I wonder how the 10% ownership at 2B valuation would be interpreted by investors. If it's viewed as a fairly liquid investment with low risk of depreciation then yeah, I could see Nvidia's strategy being quite the way to pad numbers. OTOH, the valuation could be seen as pure marketing fluff and mostly written off by the markets until regulations and profitability are firmly in place.
But there are even fewer EU VCs.
But one thing Mistral could do is have a free foundational model, and have non-free (as in beer, as in speech) "pro" models. I think they will have to.
Deploy larger, fine tuned variants and charge for them.
There’s a reason we don’t have the data set or original training scripts for mistral
It's like 1.6gb, ones coming are better and smaller https://x.com/EMostaque/status/1732912442282312099?s=20
I think the large language model paradigm is pretty much done as we move to satisficing tbh
I've been trying out all sorts of open models, and some of them are really impressive - but for my deployed web apps I'm currently sticking with OpenAI, because the performance and price I get from their API is generally much better than I can get for open models.
If Mistral offered a hosted version which didn't have any spin-up time and was price competitive with OpenAI I would be much more likely to build against their models.
I suppose they could be the Google to everyone else's Yahoo and Dogpile, but I expect that to be a hard game to play these days.
Besides, we don’t know what future opportunities will unfold for these technologies. Clearly there’s no shortage of smart investors happy to place bets on that uncertainty.
HN could really elevate the discourse if they flagged the submarine ads of VCs
They lose that fight a long time ago though. It seems they don't even try to pretend anymore.
While it may feel like a low moat if anyone can spin up a cloud instance with the same model, it's still a reasonable starting point. I think they will also be getting a lot of EU clients who can't/don't want to use US providers.
If the commerically-served model has improved capability and is exclusive to Mistral's service, there is a possible moat there.
Or, I guess look at how size, energy use and speed of computer hardware evolved over the past 70 years. Point is, implementation being, right now, "resource/energy expensive and murky at best" is how many very powerful inventions look at the beginning.
> If someone is thirsty, the water is the most important part, not the type of glass:)
Sure, except here, we're talking about one group selling a glass imbued with breakthrough nanotech, allowing it to keep the water at desired temperature indefinitely, and continuously refill itself by sucking moisture out of the air. Sometimes, the type glass may really matter, and then it's not surprising many groups strive to be able to produce it.
It is nuts to me that we have 100M computers capable of running LLMs properly, and yet only a tiny fraction of them does.
Heck, let us do p2p, and lend our computing power to others.
Let us build a personalized LLM.
This is, IMHO, a really interesting path forward. It seems no one is doing it.
No matter who wins, they’ll need those sweet GPUs and fabs.
[0]: https://news.ycombinator.com/item?id=38522873 [1]: https://news.ycombinator.com/item?id=38533725 [2]: https://news.ycombinator.com/item?id=38580758 [3]: https://news.ycombinator.com/item?id=38593526
- is what Mistral does better than Meta or OpenAI?
- will LLM become eventually open-source commodities with little room for innovation or shall we expect to see a company with a competitive advantage that will make it the new Google? in other words, how much better can we expect these LLM to be in the future? should we expect significant progress or have we reached to diminished returns (after all, this is only statistical prediction of next word, maybe there's an intrinsic limitation of this method)
- are there some sorts of benchmarks to compare all these new models?
I’m cracking up. I don’t need to be a rocket scientist to read this and immediately conclude it’s AI-generated. I mean, they didn’t even try to hide that. Haha.
https://sifted.eu/articles/ai-startup-aleph-alpha-raises-500...
The promise of LLMs is not in chatbots (imho). At scale, you will not even realize you are interacting with a language model.
It just happens to be that the first, most boring, lowest hanging fruit products that OAI, Anthropic, et al pump out are chatbots.
Alot of work for a company with no commercial offering off the bat. And possibly an insurmountable amount of work for new players trying to enter.
If you have no commercial offering it doesn't apply to you at all in the first place
Complex/simple is not really the right way to think about training these models, I'd say its more arcane. Every mistake is expensive because it takes a ton of GPU time and/or human fine tuning time. Take a look at the logbooks of some of the open source/research training runs.
So these engineers have some value as they've seen these mistakes (paid for by Meta's budget).
https://medium.com/@datadrifters/mistral-7b-beats-llama-v2-1...
Until a company is consistently showing growth in revenue and a path to sustainable profitability, valuation is essentially wild speculation.
OpenAI is wildly unprofitable right now. The revenue they make is through nice APIs.
What is Mistral’s plan for profitability?
Right now stability AI is in dumps and looking for a buyer.
Only companies I see making money in AI are those who live like cockroaches and very capital efficient. Midjourney and Comma.ai come to mind.
Very much applaud them for open release of models and weights.
Initial high valuations mean the founders get a lot of initial money giving up little stock. This can be awesome if they become strongly cash-flow positive before they run out of that much runway. But if not, they'll get crammed hard in subsequent rounds.
The more key question is: how much funding did they raise at that great valuation, and is it sufficient runway? Looks like €450 million plus an additional €120 million in convertible debt. Might be enough, depending on their expenses...
Like it takes time to make lots of money and it’s really hard to build state of the art models.
Reality is this market is huge and growing massively as it is so much more efficient to use these models than many (but not all) tasks.
At stability I told team to focus on shipping models as next year is the year for generative media where we are the leader as language models go to the edge.
To my mind they just seemed to be responding to the slightly clickbait-y title, which focuses on the valuation, which has some significance but is still pretty abstract. Still, headlines love the word "billion".
The straight-news version of the headline would probably focus more on a16z's new round.
The thing is I don’t want the pro-open-source players to fizzle out and implode because funding dried up and they have no path to self sustainability.
AGI could be 6 months away or 6 decades away.
E.g Cruise has a high probability of imploding. They raised too much and didn’t deliver. Now California has revoked their license for driverless cars.
I’m 100% sure AGI, driverless cars and amazing robots will come. Fairly convinced the ones who get us there will be the cockroaches and not the dinosaurs.
AGI is a bit of a canard imo, its not really actionable on a business sense.
But I might have a bias because I was following along as the company was built from whiteboard diagrams to what it became.
https://github.com/skorokithakis/ez-openai/
With all that money, I would have thought they'd be able to design more user-friendly APIs. Maybe they could even ask an LLM for help.
Instead of "path to profitability", I think path to ROI is more appropriate, though.
WhatsApp never had a path to profitability, but it had a clear path to ROI by building a unique and massive user base that major social networks would fight for.
Do we know some of its numbers? How many paid subscribers do they have? I pay for two subscriptions.
On profitability, For all the new comers, I don't think anyone can wager that any of them is going to make money. Capital efficiency is overrated so long as they can survive for the next year+, they are all trying to corner the market and OpenAI is the one that seems to have found a way to milk the cow for now. I truly believe that the true hitmakers are yet to enter the scene.
I think it would make much more sense to focus on the "reality side" of the transaction, e.g. "Mistral AI received a €450 million investment from top tech VC firms."
I know it can be wrong, but usually when it is, it’s obviously wrong
Support workload on our Slack was reduced by 50-75% and the output is steadily improving.
I wouldn’t want to go back tbh.
- Writing: emails, documentation, marketing - Write a bunch of unstructured skeleton of information. Add a prompt about the intended audience and a purpose. Possibly ask it to add some detail.
- Coding: Especially things like "Is there a method for this in this library" - a lot quicker than browsing through documentation. Some errors - copy-paste the error from the console, maybe a little bit for context, and quite often I get the solution.
And API based:
- Support bot
- Prompt engineering of some text models that normally would require labeling, training, and evaluation for weeks or months. A couple of use cases - unstructured text as an input + prompt, JSON as an output.
more efficient than just googling "<method description> <library name>"?
I use DALLE3 extensively for my woodworking hobby, where I ask it to come up with ideas for different pieces of furniture, and have constructed several based on those suggestions.
For work I use it to write emails, to come up with skeletons for performance reviews, look back look ahead documents, ideas for what questions to bring up during sprint reviews based on data points I provide it etc.
It’s just so much more efficient in getting the answers I need. And it makes a great pair programmer partner.
The good startups are building, fine tuning, and running models locally.