Mistral looks set to challenge AI frontrunners Google and OpenAI
ft.com
ft.com
GPT3 to 4 took almost 3 years. Google Gemini took around 3/4 of a year.
Mistral 7b to Mixtral 8x7b and simultaneously Mixtral Medium was ~82 days.
That's crazy fast.
Edit: finally archived https://archive.md/RrXWo
This leaderboard is the only one most people trust as it uses an ELO system with humans blindly comparing and rating them in the chatbot arena.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
Current leaders are OpenAI, Anthropic, Mistral, Google, 01-ai, and then various Llama finetunes and Llama itself (Meta)
It seems their vision of "alignment" does not align with the users vision of a competent assistant even remotely.
I would never use Claude for personal use. However, my employer wants to make sure nothing reputation-harming happens. That's a lot more important than model quality; everything is light years ahead of where we were four years ago. For a lot of applications which face customers / citizens / students / ..., NEVER screwing up is a lot more important than quality.
For my own use, I prefer interacting with soul, humanity, attitude, and edge.
For my employer's use, it's different.
I think there's space for both. As an investor, I'd be bullish on both Anthropic and something with no safety built in. I'd be a lot less bullish on what's in between.
However I've found that using it as a personal assistant and for some automation tasks it suffers.
The fact that it has only gotten worse as judged by humans doesn't inspire confidence for me. I could totally understand offering two options, but they're going so far it's actually unusable for many tasks.
The famous case it refused to "kill a python process".
I will pick the best tool for the job I'm doing, be that writing product descriptions, conversational agent, or tech support. Runner-up has a chance -- for example, by lowering margins, or simply by being subsidized by investors in hopes of moving into #1.
However, the difference between "unusable" and "fourth-best" is negligible in terms of business returns. I won't pick your product.
There isn't a snowball's chance of "safe AI" being #1, or even #5, for what you want to use it for. It might as well be unusable. It needs to be #1 in the niche it's targeting.
(The above isn't universal; there are places where bundling many types of functionality has synergy; this just isn't one of them).
Also shouldn't starcraft be solved for ages now? OpenAI the top teams in DoTA a long time ago, and i assumed the coordination problem would be a big issue.
(I think their bots started winning 1v1 SF mid a long time before that...)
Gemini Pro as available today already puts Google at #4 by company in the chatbot arena.
But given the amount of BS in this space these days, take with an accordingly sized grain of salt. I'll believe Google is competitive with the SotA when I see it myself.
Ahh, the magic words that cause any venture to fail
1. Doesn't have to pay the obscene alignment tax that OpenAI/etc have to (something between 20-50% of from what i understand) nor worry nearly as much about """the brand""" and the baggage which comes with that.
2. Likely will have/has the EU behind them out of pragmatic protectionism realpolitik verus the US (which basically started with GDPR) - though it might end up being a duo with Aleph Alpha.
3. Generally has the widespread support of the hobbyist/opensource community (for whatever aims those may be), both out of ethical/moral consideration and performance/quality reasons
4. Seems already *quite* competitive with GPT3. I rarely find i need to invoke GPT4 and when I have to, I'm annoyed with the latency and milquetoastian nature of the damn thing.
Wish I could purchase stock in these guys honestly
People have an axe to grind with OpenAI because they purposely align their model to be boring as the intended audience is companies embedding it in their own products and try to fight the bias inherent in a model trained on the writings of the average chronically online Redditor. The frustration isn't unfounded, if you want to use the model for programming having to pay the 'alignment tax' because someone else is using the model for a customer service bot and can't match energy with an angry customer sucks.
Instruction tuning helps a lot with this but what a lot of people mean is the refusal to do things. You get to chose how "aligned" it is, for some usecases like talking to customers you definitely want something very "safe" (won't start using slurs or something terrible). But for direct usage you generally never want it to refuse to do anything.
Checkout Anthropic on the extreme side - every iteration of Claude has gotten worse on the chatbot arena (elo based on humans blindly comparing responses).
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
The 'common' example that people reference is, for instance, making the model non-offensive means it has to spend.. something... on that directive, if you will. (which also led to absurdities such as refusing to help people with bash commands that involve having to ```kill``` a process...)
But even for the purpose of an instruct model (this is what peeves me off), making it answer questions and take instructions makes it *worse* at many creative tasks because you're constraining it's behaviour to Q&A -- though this is a long tangent...
prudish
square
no fun
it just kind of means "overly apologetic, meek, timid, unwilling to ever voice anything that may be even mildly controversial", etc. It apparently originates from H.T. Webster "The Timid Soul"