Four new models that are benchmarking near or above GPT-4
simonwillison.net
simonwillison.net
The new claude-3-opus-20240229 model got the highest score ever on this benchmark, completing 68.4% of the tasks. While Opus got the highest score, it was only a few points higher than the GPT-4 Turbo results. Given the extra costs of Opus and the slower response times, it remains to be seen which is the most practical model for daily coding use.
The team details the slight modification they made to the metric which is basically a tweak on maj@32 that worked in their favor. They choose to do the majority vote of 32 samples unless the probability of the highest sample was too low and then they chose greedy, they also had to learn what "too low" was. Sound convoluted and a bit contrived? They even show plots that clearly demonstrate they settled on the metric because it's the one where GPT-4 results remain they same, but Gemini wins.
Don't get me wrong, LLM evaluation is essential. But in order to do this you need to be very consistent between comparisons and even then the results can be misleading.
/chat Here is my message
In your code to invoke the LLM. Copilot is seamless and I've grown accustomed to how it works. Ideally I could also put in my own local LLM to test with as well. I've seen Continue.dev but it seems to require some kind of manual invoking (I'd love to be wrong).I'd like to try swapping out GPT-4 for Claude 3 Opus but I'm not super interested in changing my workflow. I manually use ChatGPT or Copilot Chat (less so) for some cases but I really like the in-line suggestions in my editor.
The rumor, with no supporting evidence, of course, is that they training corpus includes all of LibGen.
Inflection's offering looks like a scam given that the responses that it gives are very close or the same as Anthropic's Claude 3 or even being an API wrapper around the Claude API [0] and packaged as a product.
Pi is more like a toy with performative use-cases. How the company was able to raise millions of dollars out of that absolutely make no sense whatsoever other than pure VC hype.
I'm withholding judgement until more information comes out. It's a really weird situation.
Can anybody explain what is meant by "vibes" here? Is it an industry term?
A model might score off the charts but not be particularly useful for your personal style of promoting.
Everyone I know who plays with LLMs says things like
- it didn't answer my special secret benchmark question better than previous models
- it surprised me by being able to give a satisfying answer for x that I wasn't expecting
- it always seems to give answers that are [Some negative quality]
Obviously these things are super subjective but still important and useful in a way that "answers 5th grade maths questions more accurately than previous models" sometimes isn't
Competition will encourage a price race to the bottom (marginal cost = marginal revenue), but I question whether some of these companies will be able to maintain it long-term without lots and lots of venture capital.
- lmsys ranks claude 3 opus ahead of GPT4 but behind GPT4T https://x.com/lmsysorg/status/1765774296000172289 mistral large is also well behind claude 3 sonnet, and both behind gpt4, which is... good but disappointing?.
- ai2 wildbench largely agrees with this ranking https://x.com/billyuchenlin/status/1766079601154064688?s=20
- claude summarization is better, at the cost of higher hallucination https://twitter.com/swyx/status/1764805626037993853
it's also noteworthy that gemini and claude seem to have completely solved the needle in a haystack problem vs openai/mistral - i'm not sure what breakthrough is responsible but the opensource community seems to be pointing at ringattention as the main promise
> I use these models a lot
I have very few uses of AI myself, that mostly center around writing scripts in languages I'm unfamiliar with so that I don't have to spend time looking up the specific syntax. But that would not count as "a lot".
Beyond curiosity and model testing, what do people really use AI for?
Lots of quick answers to questions: my rule of thumb is that if someone who had just read the relevant Wikipedia page could answer my question, and the stakes aren't particularly high, I'll use an LLM.
Jargon decoding - I can read academic papers now!
The LLM arena shows the performance of GPT-4 compared to Claude 3 and Mistral-large. (GPT-4 is still on top)
> Not every one of these models is a clear GPT-4 beater, but every one of them is a contender.
This gave me a window to try out Gemini. I 1.0 was better than I anticipated, but still got some of the (important) details wrong. I wish I had access to 1.5.
The competition is closing in, and the frequent downtime makes it an easy choice to try the competitors.
That said, I'm still pulling for OpenAI. We need someone to challenge the old tech companies.
I hope GPT-5 is similar to Sora - ten steps forward in terms of innovation.
What are each models strengths? I know GPT4 has a strength with code.
GPT-4: https://chat.openai.com/share/117fb1ad-6361-41e2-be59-110f32...
Claude 3 Opus: https://gist.github.com/simonw/2002e2b56a97053bd9302a34e0b83...
The GPT-4 code didn't work - it was missing some async keywords.
The Claude 3 Opus code worked perfectly.