Google's Bard shows big leap on LLM performance leaderboard
twitter.com
twitter.com
I still think they ought to launch a subscription so we can see their absolute best model running in public.
it's better to let more people interact with it because this will help training the model (get more data) so it must be free to use.
I don't know about that second part - but it would make sense that google (and others) may want to use lmsys's arena to benchmark their models.
After all, Human A/B tests are far better then the current automated benchmarks.
I would like more info from lmsys as to how they're accessing these though.
From the vertex ai:
API_ENDPOINT="us-central1-aiplatform.googleapis.com"
PROJECT_ID="test00"
MODEL_ID="gemini-pro"
LOCATION_ID="us-central1"
curl \
-X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://${API_ENDPOINT}/v1/projects/${PROJECT_ID}/locations/${LOCATION_ID}/publishers/google/models/${MODEL_ID}:streamGenerateContent" -d '@request.json'
and from the makersuite: curl \
-X POST https://generativelanguage.googleapis.com/v1beta/models/gemini-pro:generateContent?key=${API_KEY} \
-H 'Content-Type: application/json' \
-d '@request.json'Everyone else needs to pay nvidia margins.
Training is murkier as it’s more about the total performance and scalability of the system.
Does anyone know if the GPT-4 Turbo version used on the leaderboard has access to web search? I always assumed it did not, but now it doesn't seem like an apples-to-apples comparison.
https://x.com/asadovsky/status/1750983142041911412?s=20
Edit: I used the "Direct Chat" feature on lmsys to ask Bard and GPT-4 Turbo "What is the current price of Bitcoin?". Sure enough GPT-4 Turbo said it can't browse the Internet and Bard gave a real time answer from Google Search. This means GPT-4 outperforms Bard overall even without the ability to browse the web at all. Pretty impressive.
These seem like different categories; one is a model and one is a system with a model plus tools. I think it is useful to compare them, since there is a real difference in user experience. However, they ought to be prominently marked as different categories. And the lmsys guys ought to put a ChatGPT model on the leaderboard with its own search integration enabled, for a fairer comparison. And it would be cool to have other LLM+tools entries like Perplexity, Phind, etc.
> Bard, powered by the Gemini Pro-scale model, debuts at the #2 position on the independent lmsys leaderboard.
According to Jeff Dean's tweet, it looks like they have a new "Gemini Pro-scale model" being rolled out, not sure what it means by "Pro-scale" though. Also not sure if everyone already got it...
My own experience doesn't really match the arena's leaderboard. I find myself using Claude 2 much more than GPT 4/Turbo. I prefer its out-of-the-box response style, and its answers for my queries seem just as good, if not better.
Interesting, Kagi (where I use the chatbots) ranks Claude 1 as equal to (and faster than) GPT 4 (non-Turbo), and marks Claude 2's quality as on-par with 4 Turbo (albeit slower). Worth noting it's a simple 1-4-star ranking, unlike the arena's ELO numbers.
I personally use bard for general internet search stuff, because I like how google has set up the linking (and the UI). I use GPT-* for technical stuff (eg GitHub codepilot) because I find it can synthesize code better.
[0] https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WB...
Moderation means no sex, hate, illegal things and religion. I tried to talk about Allah being merciful and got my ass moderated away. I am Buddhist and when I talked about reincarnation being misunderstood that was fine. So the limits are not clear.
For example it might tell me to read some documentation and find the answer myself. So I'm thinking "Yes well, what do I need you for then?"
A: Certainly! Here is a function that makes a couple of text edits and a button.
function renderForm() {
var edit = document.createElement('input');
// The rest of the code is similar.
}I had a colleague complain to me the other day that he'd asked GPT to help him with a trivial reformatting task and it had said something like "I can't do that but I can guide you on how to do it".
agreed. that's pretty "lazy"
I gave Bard a go, after seeing Jeff Dean's tweet.
It's just as frustrating as it was, compared to GPT-4. It's simply off the question and unable to realize it's off.
I asked it to generate a chart and 3 times it came back with "here's a chart" with no chart, finally saying it doesn't have that ability.
Also I felt I keep getting more personalized results, i.e. models are somehow biased towards user. I heard they plan it, I don't know it's launched, but I feel it.
And there's also fine-tuning in the other direction - my brain got used to ways of interacting with GPT. Same as with Google, I just somehow subconsciously know how to write prompts that get me what I want.
If this aint anecdotal evidence, I dont know what is. You need to assume a whole lot to draw these conclusions from prompt output comparison.
Heck, just compare a couple outcomes from ChatGPT and you should conclude the same thing, assuming rationality of course.
Also what a wild comparison, because afaik chat gpt can’t make charts either.
Wait so if Google suddenly started charging for Bard, it would be instantly better?
They can spend billions giving away compute (cost) and be extremely efficient (high utilization) in doing so.
It's either to expensive or they just don't think it's that important to have a competitive, accessible model out there (which is a valid stance imo).
It will write Python code to use matplotlib and the resulting chart will appear in the chat. I'd show an example but their sharing feature doesn't work with images for some reason.
I tried this on the free version you can get to without paying anything, I didn't realize the advanced version supported this.
Here you go https://imgur.com/a/MT04viM
Disclaimers: I don’t work for Google
Getting close, but this will force OpenAI to come out with gpt5 for another big round of catch up
"mommy, can we get GPT?" "we have GPT at home darling"
That being said Chatbot Arena is a pretty wide variety of test scenarios. If fine tuning made the model perfect, all the small models would get similar scores to GPT4, which they don't. Essentially it ranks how people believe a ChatBot should respond, rather than just zero shot, 1 shot and COT type benchmarks.
You could ask a question on lmsys, check your server logs for the generated response, go back to lmsys and pick the response that your model generated.
Maybe you could also use a better model for requests from lmsys. E. g. use an unquantized model, disable censorship, etc.
I doubt any of the big players are doing that, but you never know.
But yeah we tried some things.
The interesting thing is it still has a 128k context length. This is awesome because GPT became way more useful to me once it reached this level of context.