Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
arxiv.org
arxiv.org
To get to this, I asked him to predict the return value of some simple python function (for """def func(a,b): return a*3 + 2*b)""" what's func(2)).
without requiring the explanation, the answers were not only wrong but also changed between attempts for the same input.
Asking the model to provide an explanation, however, resulted in it giving the correct answer each time.
This method still allows for improvement in the quality of answers in more complex problems, although at a bit lower correctness rate.
Other models provide better answers if you tell them it’s important for your career or that you’ll tip them $200 if they’re right, or if you tell them that it’s allowed to not know… the preamble to get max performance gets long.
The reason this works is it puts the reasoning into the context window. Transformers pay attention to everything in the context window and having a rationale sitting there - even if the model generates the rationale - helps the model do a better job of issuing the correct multiple choice answer.
It seems reasonable to me to assume it’s somewhere in the neighborhood of 20 billion, but I agree it is worthwhile to recognize that we don’t actually know.
Even so, Google treats the Gemini Pro Vision model as a separate model from Gemini Pro, so it could have separate parameters that are dedicated to vision (like CogVLM does), and that wouldn’t impact the size of the model as far as text-tasks are concerned.
My experience with Gemini pro is that it often completely and spectacularly misunderstands the ask. No other GPT comes close to getting it that wrong. As a result, I don't have a high level of confidence in it.
I find Mixtral to be quite impressive. It is by far the fastest and gives me good output. At least for big documents kind of work.
It just added support for configuring an arbitrary number of custom endpoints for your locally hosted models, and is improving a ton.
Then, when they do get over the 3.5 line, they will just graveyard it anyway.
They've fully earned a reputation as a marketing parasite.
This is the impression I have of google as well. Especially in the AI space.
Remember that “AI” scheduling a haircut or something that google presented in 2017?
That doesn’t mean you’re any good at building cars.
See https://storage.googleapis.com/deepmind-media/gemini/gemini_...
By the way, I linked to this report when it came out, but was labeled a "dupe" because someone else had already posted a link to the flashy website without any technical details.
I cannot believe someone in Google approved that video but then again I cannot believe that a company would be launching 10 different chat apps.
Honestly Gemini Pro is so unimpressive I expect even GPT-3.5 will beat Gemini Ultra in terms of utility, at least for the next 6 months.
I couldn't glean from the post; does anyone know what the test suite is?
I notice failure in common sense reasoning quite often in Bard and ChatGPT and while I haven't done a systematic analysis of the categories of mistakes, some patterns I've seen include: lack of self-reflection, not noticing internal inconsistency in output, lack of awareness of what would be "common" common sense, etc.