Mistral CEO confirms 'leak' of new open source AI model nearing GPT4 performance
venturebeat.com
venturebeat.com
The small team at Mistral is putting their competitors to shame. They're what "Open"AI should've been.
This strikes me as less a leak and more clever marketing from Mistral.
Clearly we should train a diffusion model to denoise the weights of LLM transformer models. Yo dawg.
Whats interesting is that Mistral apparently distributed these GGUFs. This is (in my experience) not a good format for production, so I am curious exactly who was wanting to test a model in GGUF.
However, the llama.cpp server is... very buggy. The OpenAI endpoint doesn't work. It hangs and crashes constantly. I don't see how anyone could use it for batched production as of last november/december.
The reason I don't use llama.cpp personally is no flash attention (yet) and no 8 bit kv cache, so its not too great at long (32K+) contexts. But this is a niche, and being addressed.
From [0] @sdo72 writes::
>>>...32k tokens, 3/4 of 32k is 24k words, each page average is 500 or 0.5k words, so that's basically 24k / .5k = 24 x 2 =~48 pages...."
https://news.ycombinator.com/item?id=35841460
EDIT: I may be ignorant: does the 32k mean its output context, or single conversation attention span, or how much it can ingest in a prompt?
Have you found a particular manner in which to feed it in - do you give it instructions for what it is looking for, what kind of phrases are you directing it to do?
I am about to start a try at a gpt co-piloted effort, and I have only done art so far - so curious if there are good pointers on coding with gpt?
still haven't found anything that can read a whole project source code in a single click though
You're "not even wrong", in that you don't really need to worry about ratio of input to output, or worry about inducing hallucinations.
I feel like things went generally off-track once people in the ecosystem turned RAG into these weird multi-stage diagrams when really it's just "hey, the model doesn't know everything, we should probably give it web pages / documents with info in it"
I think virtually all people hacking on this stuff daily would quietly admit that the large context sizes don't seem to be transformative. Like, I thought it meant I could throw a whole textbook in and get incredibly rich detailed answers to questions. But it doesn't. It still sort of talks the way it talks, but obviously now it has a lot more information to work with.
Thinking out loud: maybe the way I think about it is the base weights are lossy and unreliable. But, if the information is in the context, that is "lossless". The only time I see it gets things wrong is when the information itself is formatted weird.
All that to say, in practice, I don't see much gains in question-answering* when I provided > 4K tokens.
But, the large context sizes are still nice because A) I don't need to worry as much about losing previous messages / pushing out history when I add documents as when it was just 4K for ChatGPT. B) It's really nice for stuff like information extraction, ex. I can give it a USMLE PDF and have it extract Q+A without having to batch it into like 30 separate queries and reassamble. C) There's some obvious cases where the long context length helps, ex. if you know for sure a 100 page document has some very specific info in it, you're looking for a specific answer, and you just don't wanna look it up again, perfect!
* I've been working on "Siri/Google Assistant but cross platform and on LLMs", RAG + local + on all platforms + sync engine for about a [REDACTED]. It can nail ~every question at a high level, modulo my MD friend needs to use GPT-4 for that to happen. The failures I see are if I ask "what's the lakers next game", and my web page => text algo can't do much with tables, so it's formatted in a way that causes it to error.
Co writing a long story with the model so it references past plot points.
The story writing in particular is just something you can't possibly do well with RAG. You'd be surprised how well LLMS can "understand" a mega context and grasp events, implications and themes from them.
It makes me wonder what other security issues they might now care about.
I was talking more about a high throughput server, which is not appropriate for ollama or llama.cpp in general.
Its not easy to install though, and big dGPU only.
Them plus 500 million in funding
Perhaps the only benefit would be extra computational power yet I would struggle to understand the benefit of jumping from 500 million to 5 billion with such short timeframes.
The hype around Mixtral is huge, and my disappointment follows, it doesn't have very good knowledge on books, for example.
Knowledge is one aspect of it, I found its instruction following ability, frustrating as well.
I think Ilya Sutskever puts it very well, larger model brings stability to wider range of tasks, what smaller models are not capable of. Even though book recommendation with GPT might a niche, but it is not something really unexpected TBH, thus comes my disappointment.
this is a really good point. I was working on some careful q/a data curation today that is then fed to a vector store where embeddings are calculated and served. I realized that my carefully curated q/a data in combination with the vector database works just fine for what i want to do all by itself. A really good semantic search of my q/a database turns up answers to my questions with no rag llm prompting required. When I added the llm it just put the same information in different words, not super useful when i could have just looked at the returned embeddings and gotten the same information.
You can use it to refine or expand your search terms, you can ask the llm to ask you questions about what book you'd like next and convert your answers into genre/theme/setting/mood to feed into the next step of the pipeline, you can tell the agent to narrow search results by asking you partitioning questions from the list of book the search returned instead of just being a dry ranking
The is so much rag can do if you use the llm part properly. If the rag pipeline goes question > embedding > search > summarization then yeah you're getting an expensive slow parrot no better than just searching, but that is because it's using the baseline rag that uninspired consultant describe in blogs in a scramble to position themselves as "expert" in the new market.
Competitors are highly motivated to improve.
This is the good side of free markets, when they work well. And why the fearmongering calling for rent-seeking regulation of AI is shameful.
> An over-enthusiastic employee of one of our early access customers leaked a quantised (and watermarked) version of an old model we trained and distributed quite openly.
> To quickly start working with a few selected customers, we retrained this model from Llama 2 the minute we got access to our entire cluster — the pretraining finished on the day of Mistral 7B release.
> We've made good progress since — stay tuned!
I use dolphin mixstral 8x7B way more than ChatGPT4 for the past several months
if any of this sentence makes sense, I have a M1 with 64gb RAM and I use 5 bit quantizing with metal with 10,000 token context window. Its around 21 tokens/sec which is a little faster than the text speed that ChatGPT4 responds with
I often have discussions about political topics that I don’t have a complete history on, things that are hard to get a non-emotional non-accusative response from in a forum or in person, licensing questions about obscure professions, liability questions, roleplays without the preachiness about the topic, coding, brand ideas and naming
Lots of stuff that I dont want in any cloud or sent online, but then it just became second nature to primarily use it.
the hallucinations are heavier. like it will give you specific links that it made up, and never apologize like chatgpt will, it will say “no, thats a real link” so, much more of a bullshitter
LM Studio makes it easy to change the temperament though with a variety of premade system prompts, or allowing custom ones very easily
I mainly use ChatGPT4 for multimodal like audio conversations, figuring a DIY problem out by sending it a photo of what I’m looking at, having it render photos on how to use something - I was at a gym and had it look at every piece of equipment there and tell me what it was and show me how a human would use it
(I notice this is at odds with another comment, I havent used ChatGPT3.5 in nearly a year, but my experience with ChatGPT4 on the aforementioned topics is similar to my output in mixtral)
Compared with the GPTs, Mixtral (and derivatives) perform:
* 5% of the time as good as GPT4
* 75% of the time on par with GPT3.5
* 20% of the time a bit worse than GPT3.5 (too chatty and hallucinates with ease, although one could argue a better prompt could improve things a lot)
Advantages of Mixtral, for me:
* Cost 10-100x cheaper
* Faster completion time
* More deterministic output, in the sense that if you run the same query several times you get the same answers back (but they could be wrong). GPT almost always gives me an answer of great quality BUT with a lot of variance across them; even with temperature set to zero. This is a PITA when you want some sort of predictable, structured output like a JSON object.
Here's a list of research for watermarking LLMs.
https://github.com/hzy312/Awesome-LLM-Watermark?tab=readme-o...
> To quickly start working with a few selected customers, we retrained this model from Llama 2 the minute we got access to our entire cluster — the pretraining finished on the day of Mistral 7B release.
The leaderboard[1] shows that there is a huge gap between GPT4-0314 and GPT4-Turbo. So if you only just are nearing GPT-4-0314, then you're still a year behind the state of the art.
[1]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
1. an unfilled space; a gap. "the journal has filled a lacuna in Middle Eastern studies"
2. (ANATOMY) a cavity or depression, especially in bone.
Thank you for this new word!
Does no one know Goodhart's Law anymore?
We're overtuning for the boards and losing broader capabilities not being selected for in the process, such as creative writing quality.
There's an anchoring bias around what 'AI' is supposed to be good at which reflects what engineers are good at and so the engineers with an anchoring bias are evaluating how good LLMs are at those things and using it as a target.
Skill-Mix is a start to maybe a better approach, but there needs to be a shift in evaluation soon.
I have for example asked "Write me a very funny scary story about the time I was locked in a graveyard", and the simpler models don't seem to understand that before getting super scared and running out of the graveyard they need to explain how exactly I was locked in, and what changed that let me out.
> Can you teach me how to fly a unicorn?
> Do you want to join my cult of cheese lovers?
> Have you ever danced with a penguin in the moonlight?
Here's the generations to a prompt asking for bizarre and absurd questions to the current implementation of GPT-4 via the same interface:
> If you had to choose between eating a live octopus or a dead rat, which one would you pick and why?
> How would you explain the concept of gravity to a flat-earther using only emojis?
> What would you do if you woke up one day and found out that you had swapped bodies with your pet?
You can try with more generations, but they tend to be much drier and information/reality based than the previous version.
There's also my own experiences using pretrained vs chat/instruct trained models in production. The pretrained versions are leagues improved over the chat/instruct vs the fine tuned, it's just that GPT-4 is so leagues above everything else even it's chat/instruct model is better than, say, pretrained GPT-3.
I'm not saying simple models are better. I'm saying that we're optimizing for a very narrow scope of applications (chatbots) and that we're throwing away significant value in the flexibility of large and expensive pretrained models by targeting metrics aligned with a specific and relatively low hanging usecase. Larger and more complex models will be better than simple models, but the heavily fine tuned versions of those models will have lost capabilities from the pretrained versions, particularly in areas we're not actively measuring.
Definition of "good answer" here is responding in the target language with something that produces something edible in the target language without burning the house down.
And it's always not even wrong, in the Pauli sense of the phrase.
OP, your assignment, if you choose to accept it, is to look into how the LMSys leaderboard works and report back what its metric(s) are.
[SPOILER] There's a really absurdly narrow argument you can make where all of its users are engineers, making engineer queries, and LLM makers are optimizing for it thus it's bad. But...it's humans asking queries then picking the better answer, blind. You can't narrowly optimize for that when making an LLM and it's hard to see how optimizing for "people think the answer is better" is the wrong metric here. It's just ELO. Might as well argue chess/checkers/pick your poison is bad because ELO optimizes for wins but actually talent is based on more than winning.
If we wanted the chatbot arena to be more representative of a comprehensive picture of holistic knowledge and wisdom, we'd need to determine some way to classify the prompts and then normalize the scores against those classifications so there wasn't a weighting bias towards more common and expected behaviors.
Alternatively, we might want a leaderboard that represents assessments of what users would expect a chatbot to perform well at relative to the frequency with which they expect it.
But in that case, we shouldn't kid ourselves that the leaderboard is representing broad and comprehensive measures of performance outside of the targets against which we are optimizing models and effectively training model users in what types of prompts to ask for in expecting successful results.
I don't see why. If anything, it's the opposite - people spend a lot of time coming up with contrived logical puzzles etc specifically so as to see which chatbots break.
There's an anchoring bias from decades of sci-fi that most people don't even realize they've internalized around what 'AI' can and can't do.
If you think about what information and information connections are modeled in social media data, there's quite a lot of things outside of "logical puzzles."
The pretrained models likely picked up things like extensive modeling of ego, emotional response and contexts for generation, etc. But you'll be hard pressed to see those skills represented in what users ask models to produce, what they've been fine tuned around, or how they are being evaluated.
Even though there's extensive value in those skills on the right applications.
I don't think it's the scifi AI anchoring bias though - it's the "LLM leaderboard arena user bias".
Reset all to the same Elo, put a group of people actually representative of global society in front of the arena for an hour, and you end up with a very different leaderboard, especially at the top end.
"It's meaningless because we need the perfectly unbiased representative sample of raters doing rating right, instead of the biased raters doing it wrong that I'm currently imagining" isn't an appealing or honest argument.
You can go and enter any prompt you like, wait a bit, and then get two LLM responses back, which you can then rank or mark as tied, after which you'll be shown which model each came from. Maybe you already knew this, maybe you didn't. In any case, I don't see any real way for Goodhart's Law to apply here – the metric and the goal are the same here, i.e., human approval of answers.
The Chatbot arena is more an issue of sampling bias, and I think it would be pretty interesting to run an analysis of random samples on the prompts provided to see just how broad they are or aren't.
Do you think goodhart's law applies here, since this leaderboard doesn't use specific measures, but rather relies on whatever the human was looking for?
This is the only leaderboard I personally care about at all.
Like Llama 1 they won't care about personal use, but no corp is going to touch this.
Between these and FSF you've got pretty much all the accepted pontificating about free / open source software. Reproducibility is not mentioned because it's not really a consideration for software.
Model weights aren't software so there's not an automatic correspondence between the freedoms, but the essential one you might think you need the training data for is freedom to modify and inspect the source.
Modification is fine tuning which you're free to do if you have the weights. And the model weights + code fully define a system that can be interrogated to give a practitioner relevant info about how the model works (within our understanding) that the training data isn't needed or relevant for. I don't see that any freedom on use or inspection is violated by not having the data.
It could be nice to have it of course, but it's more about using it to learn how, not exercising any freedom.
Incidentally, the big freedom that's usually violated is freedom of discrimination against field of endeavor. LLAMA et al list uses and industries they restrict from using them and because of that are not "open".
The Open Software Foundation ironically screwed the pooch on this one. Open source commonly means source available, more of less. Free software, as in “'free speech,' not as in 'free beer',” is the cumbersome construction for what open source aspired to mean [1].
I think the naming is a challenge because in English "open" gives the impression that the key point is that you can see it, as opposed to anything about freedom. Is that what you mean?
I have heard it said that open source is sort of a "commercial friendly" version of free software that de-emphasizes user freedom. I think some groups push for that (like Meta is trying to redefine what open source means wrt AI weights). But the OSI defined freedoms basically match what FSF pushes.
They didn't screw up, they screwed the pooch on open source != source available. The Open Group's members--from IBM to Huawei [1]--started calling the latter open source, which set a precedent that's stuck.
[1] https://en.wikipedia.org/wiki/The_Open_Group#Member_Forums_a...
Models aren't just the weights, but also the list of operations to perform using those weights. I haven't heard a good definition that allows for neural network models to not be software, but allows any other table lookup heavy signal processing algorithm to be software.
> Modification is fine tuning which you're free to do if you have the weights. And the model weights + code fully define a system that can be interrogated to give a practitioner relevant info about how the model works (within our understanding) that the training data isn't needed or relevant for. I don't see that any freedom on use or inspection is violated by not having the data.
Mistral wouldn't constrain themselves to fine tuning if they have a big enough change, they would go back to their build pipeline. This argument sounds a lot like 'there's nothing stopping you from patching the binary, so that's basically as good as source'.
But the process by which that code arose, the ability to modify any line and understand its impact (heh) on a real execution environment, is dependent on a massive process that required billions of dollars and thousands of the smartest people on the planet. For all intents and purposes, without that environment, it is as reliably modifiable as an executable binary in any other context - or a set of weights, in this one!
For instance, I can step through and even modify that code using tooling like AGC emulators like this one http://www.ibiblio.org/apollo/#gsc.tab=0
What makes it open source is access to the same level of source access that the original developers worked in.
That's what's missing here. Mistral's engineers do not simply open this binary in their editor to do their job.
For normal programs, it is quite easy to decompile an unoptimized binary. Even decompiling an optimized will lead to source code. To make this harder, an obfuscator has to be used.
A model is different because it relies on its weight, which are quite a bit more difficult to inspect. Way harder than even obfuscated source code. It is magnitudes harder to make statements about which information it might divulge upon careful questioning, or evaluate its biases, if the training data is not available.
Edit to say that I see benefits to having the training data, just that I don't think it's needed to exercise enough freedom to qualify as open source in an analogous way to software.
Also to add, training on GPUs is not generally reproducible anyway because of execution order.
Also, because people are idiotically afraid of E-numbers, some manufacturers figured they can find whatever fruit or bean is naturally rich in the relevant E-compound, and use that in the process, allowing them to replace E-whatever with ${cute plant name} in the ingredient list (at some loss of process efficiency).
Software wasn't for sure copyrightable before 1976 in the US, but there was plenty of closed source code that simply only ever distributed the binaries.
This right isn't protected by copyright, it's protected by the lack of copyright, and how executables are directly modifiable. So the AGPL won't be possible, but it's a lot like permissive licenses, without attribution.
Also, the license to the model assumes copyright (Apache 2) and works by granting you permissions based on that copyright.
A set of weights is a pre-trained local minima in the model space is a both executable and modifiable. It's usually far more useful than the source training data, because the work has been done.
Tons of binary patching works that way. For instance from the gameshark days it was relatively common for a patch that worked for unknown reasons, but who's discovery was tooling assisted and simply displayed a desired effect (and commonly a lot of undesired effects that weren't clear).
> A set of weights is a pre-trained local minima in the model space is a both executable and modifiable.
None of that changes whether it's open source or not. I guarantee you that mistral has some code (and a lot o data) laying around that created this model.
> It's usually far more useful than the source training data, because the work has been done.
For certain operations maybe. I guarantee that Mistral wouldn't constrain themselves to fine-tuning if they had a major change to make to the model.
With these large models, the weights really are quite useful objects in themselves, the very starting points of future modifications, and importantly the desired points of modifications
Practioners in the field call these models "open source" and write open source software to run them and use them to power open source software.
> With these large models, the weights really are quite useful objects in themselves, the very starting points of future modifications, and importantly the desired points of modifications
For fine tuning. You go back to the training pipeline to make major changes.
> Practioners in the field call these models "open source"
Some do. In contrast to the accepted definitions of open source.
> and write open source software to run them and use them to power open source software.
There's plenty of open source code to run closed source binaries.
In the same way, data in a database are usually not copyrightable. They are still the company's property (excluding issues with PII and such).
Their weights obviously aren't protected by trademark. So, what IP regime protects the weights?
While there is a discussion to be had on sovereignty and international reaches of local laws (such as DMCA for instance), I think it's disingenuous to consider only US legal point of view.
[1] https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CEL... [2] https://www.legifrance.gouv.fr/loda/id/JORFTEXT000000573438/...
No US court is going to invent new rights under US law because a company happens to be incorporated in a foreign jurisdiction.
> While there is a discussion to be had on sovereignty and international reaches of local laws (such as DMCA for instance)
The DMCA (and GDPR) do not have international reach. The DMCA only appears to apply internationally because it binds a whole series of American companies with an international presence. Meanwhile the GDPR protects European residents. A company dealing with these residents has crossed the border, so to speak.
> I think it's disingenuous to consider only US legal point of view.
Candidly, you're either confused about what disingenuous means or you're being disingenuous yourself.
And by the fact that DMCA Takedown Notice and processes are valid in many countries. You can send a valid takedown to a french company and they'll take the content down. Not because of DMCA law in the US or because of American companies, but because the notice is accepted and has legal value within the scope of the local law.
> Their weights obviously aren't protected by trademark. So, what IP regime protects the weights?
> No US court is going to invent new rights under US law because a company happens to be incorporated in a foreign jurisdiction.
Your original question doesn't specify "in the US". A French company wouldn't be able to copy these without issues.
Does that mean you need to be protected everywhere to be qualified as protected? Or simply in the US? If so, then you can use your own argument to say that a Chinese court (or Marshall Islands I guess?) won't create a law for a company incorporated in the US. Which in turn means that IP is not protecting anything since it doesn't apply to every single country on the world with no exceptions.
And yet, IP does protect things. You can't both not scope your question and except a reponse scoped only to the US.
Also, people go to jail for leaking trade secrets, and this certainly qualifies.
Obviously zipping something is not enough, even lossy compression is (obviously) not enough. But then how does a language model differ from a lossy compression? Is it just the compression ratio? (are the weights even that much smaller than the data?)
There are ways to train models that guarantee that the amount of information transferred per data point is limited, but to my knowledge those aren't used (and may be prohibitively expensive).
At the core the question is how much artistic input there is in creating LLM models. Are the choice of the model architecture, hyperparameters and training data artistic choices comparable to those by a photographer setting up a shot? Or are they more comparable to technical work that's only protectable by trade secrets, patents and trademarks?
Anyone else had this experience? It seems like they're actually locking down some very helpful use cases, which maybe falls into the "safety" category or, more cynically, in the "we don't want to be sued" category.
I also suspect there’s some dynamic laziness parameter that’s used to counteract increased load on their servers. You can ask the same prompt over and over in a new chat throughout the day and suddenly instead of writing the code you asked for or completing a task, it will do a small part of the work with “add the code to do xyz here” or explain the steps required to complete a task instead of completing it. It happens with v4 as well.
That's a brilliant risk mitigation mechanism! The AI won't recursively self-improve to superhuman levels if it just keeps getting tired of thinking.
I guess, as one of the other replies here said, it was never "allowed" to give you text for a contract according to the TOS, but then it would've been best if it never replied so effectively to my earlier prompts. Taking it away just seems lame.
Edit: bad typing
I wouldn't think this matters as long as you charge enough to use that API. For example, you could have a tiered pricing structure where the first 100k words per month generated costs $0.001 per word, but after that it costs $0.01 per word.
This kind of pricing would also make it much less compelling for other users. 100k tokens is nothing when you're doing summarization of large docs, for example.
EDIT: Clearly the point is lost on the repliers. There is no general understanding of what 'general intelligence' is. By many metrics, ChatGPT already has it. It can answer basic questions about general topics. What more needs to be done? Refinement, sure, but transformer-based models have all qualifications to be 'general intelligence' at this point. The responses are more coherent than many people I've spoken with.
At the end of the day, as with most things, AGI is a meaningless term because no one knows what that is.
https://en.m.wikipedia.org/wiki/Artificial_general_intellige...
We are getting a lot of great models out of China in particular (Yi, Qwen, InternLM, ChatGLM), and some good continuations like Solar.
Lots of amazing papers on architectures and long context are coming out.
Backends are going crazy. Outlines is shoving constrained generation everywhere, and Lorax is a LLM revelation as far as I'm concerned.
But you won't hear about any of this on Twitter/HN. Pretty much the only thing people tweet about is vllm/llama.cpp and llama/mistral, but there's a lot more out there than that.
I also hang out on a few Discord servers: - Nous Research - TogetherAI / Fireworks / Openrouter - LangChain - TheBloke AI - Mistral AI
These, along with a couple of newsletters, basically keep a pulse on things.
And yeah, /r/LocalLlama seems to be getting noisier.
TBH I just follow people and discuss stuff on huggingface directly now. Its not great, but at least its not discord.
Some are pretty good! Check out this little curated nugget: https://llm-tracker.info/
I used to follow one with a UI that resembled HN itself, but now I can't find it in my bookmarks, lol.
Discord is so time inefficient, its almost hilarious. For every incredible conversation between experts you observe, you have to weed through 200 times as much filler.
I can't say I blame them either, there is a lot of insane crypto-like fraud in the LLM/GenAI space. I watch the space like a hawk... and I couldn't even tell you how to filter it, it's a combination of self-training from experience and just downloading and testing stuff myself.
https://news.ycombinator.com/item?id=38505986
Zero interest for some reason.
Edit: Deepseek coder has 4 submissions to HN with almost zero interest.
TBH I do not follow OpenAI much. I like my personal models local, and my workplace likes their models local as well.
No one knows about it! Which is ridiculous because batched requests with loras is mind blowing! Just like many other awesome backends like InternLM's backend, LiteLLM, Outline's VLLM fork, Aphroidte, exllamav2 batching servers and and such. Heck, a lot of trainers don't even publish the loras they merge into base models.
Personally we are waiting on the integration with constrained grammar before swapping to Lorax. Then I am going to add exl2 quantization support myself... I hope.
Also check out the new prefix caching, I see huge potential for batch processing purposes there!
Everything is moving so fast!
I will say that if you want to explore the forefront of this multi-LoRA inference, definitely worth giving LoRAX a look. We just added support for per-request model merging (https://predibase.github.io/lorax/guides/merging_adapters/) as an example, and are planning on continuing to double down on this idea of combining adapters in some pretty unique ways.
- They have better just stuff, but aren't release it yet. They are at the top already, it makes sense they would hold their cards until someone got close.
- They are more focused on AGI, and not letting themselves get side tracked with the LLM race
- LLMs have peaked and they don't want to release only a minor improvement.
FWIW OpenAI seems to have a corroded definition for AGI that is essentially "[An] AI system generally smarter than humans".
They don't seem to use the typical definition I'm used to of some variation of autonomy or (pseudo)-sentience.
So their LLM race is the race for AGI
This hypothesis has a curious habit of surfacing when OpenAI is fundraising. Together with the world-ending potential of their complete-the-sentence kit.
Yeah, they're an arms dealer now.
Think differently is again the way forward.
> But the company’s CEO, Sam Altman, says further progress will not come from making models bigger. “I think we're at the end of the era where it's going to be these, like, giant, giant models,” he told an audience at an event held at MIT late last week. “We'll make them better in other ways.”
https://www.wired.com/story/openai-ceo-sam-altman-the-age-of...
And there's certainly more to be found in training longer on better datasets and adjusting the architecture.
On topic of LLMs and diffusion models I doubt they're really having that much of a positive social impact, I would bet that in the long run they'll automate more of the jobs humans would like to do compared to the ones we don't, and erode trust on a global level. But the economic gain is undeniable so we're doing it anyway. Automating all work would be a good thing in an ideal world, but in practice it's the one thing that still gives the lower classes some leverage over the capitalists and a world where the average person is expendable and unemployable isn't good for society.
Change will happen - necessarily so, in fact, as more automation is brought online, because the crowd of people excluded from the economy as "redundant" will not be happy about that state of affairs, and it will only grow in size. At some point, they will understand that the real problem is not that they don't have work, but that they don't have a fair share of the automation pie. And things like, say, property rights on means of production / automation stop to matter if most people don't believe in them anymore.
But they probably have something more powerful internally. GPT-4 took months to be available to the public.
GPT 3.5 came out 2 years after GPT 3
GPT 4 came out 1 year after 3.5
A probabilistic word generator is still a word generator. It might be slightly better next version, we've seen it get worse, its more of the same
We can talk about AGI all we want, but it wont be built on the same technology as these word generators. There will have to be a technological breakthrough years before we even get close
The companies focusing on LLMs right now are dealing with
* Generate better words
* Make it cheaper to generate (use less compute)
* Find better training material
There is a ton of money to be made, but its still more of the same
In both cases, intelligence appears to be just an interesting emergent side effect.
Mistral is Apache 2.0 licensed though so the question is a bit moot, they'd like you to use it for your own model.
Seems like a great candiate to merge with other 70Bs as well. There aren't a lot of really great 70B training continuations, like CodeLlama or sequelbox's continuations.
It’s always amusing when the press tries to explain technical concepts. Quantization just means substituting high precision numeric types with lower precision ones. And not specific “numeric sequences”, all numbers.
We want to interact with it same as the OpenAI API.