Gemma 2: Improving Open Language Models at a Practical Size [pdf]
storage.googleapis.com
storage.googleapis.com
In my experience, smaller models tend to do well on benchmarks and fail at generalization. Phi-2 comes to mind.
It's literally the first omni-translation tool that actually works that you can run offline at home. I'm amazed that Google mentioned absolutely nothing about this in their paper.
So I guess Gemma 2 is going to become Gemini 2.0 in their truly large and closed variants then? Or is it the open version of Gemini 1.5?
prompts:
- 'Answer this coding problem in Python: {{ask}}'
providers:
- ollama:chat:gemma2:9b
- ollama:chat:llama3:8b
tests:
- vars:
ask: function to find the nth fibonacci number
- vars:
ask: calculate pi to the nth digit
- # ...
One small thing I've always appreciated about Gemma is that it doesn't include a "Sure, I can help you" preamble. It just gets right into the code, and follows it with an explanation. The training seems to emphasize response structure and ease of comprehension.Also, best to run evals that don't rely on rote memorization of public code... so please substitute with your personal tests :)
OTOH, smaller model did it perfectly.
A larger context window really helps on RAG tasks, it's frustrating that a lot of the foundational models have such small windows.
However, not all models (especially newer ones) respond well to this, which makes sense. We're working on changing the behavior in Ollama's API to be more similar to OpenAI, Anthropic and similar APIs so that when the context limit is hit, the API returns a "limit" finish/done reason. Hope this is helpful!
It's 15 ELO under Llama-3-70B on english hard prompts and 41 ELO under Llama-3-70B (the latter is actually stat sig) for general English.
Even llama-8b did better in some of my tests than Gemma 27b.
(Relevant sections of the paper highlighted.)
It was not trained or RLHFd on Arena replies or user preferences.
At the end of the day, by optimizing for leaderboard scoring, it makes the leaderboard ranking less useful as a benchmark (Goodhart's law strikes again). The Gemma team obviously isn't the only one doing it, but it's important to be clear-eyed about the consequences.
Opinions are our own and not of Google DeepMind.
Still no 27B 4-bit GGUF quants on HF yet!
I'm monitoring this search: https://huggingface.co/models?library=gguf&sort=trending&sea...
How much is pre-training dataset changes, how much is tuning?
How do you think about this problem, how do you solve it?
Seems tricky to me.
Literature has identified self-proliferation as dangerous capability of models, and details about how to define it and example of form it can take have been openly discussed by GDM (https://arxiv.org/pdf/2403.13793).
Current Gemma 2 models' success rate to end-to-end challenges is null (0 out 10), so the capabilities to perform such tasks are currently limited.
Or does the model have to later be finetuned, to not be good at certain tasks?
Or are we not at that stage yet?
Is something like tree-of-thought used, to get the best of the models for these tasks?
You can try 27b at www.aistudio,google.com. Send in your favorite prompts, and we hope you like the responses.
I don't know who writes Google's documentation or does the copyediting for their console, but it is hard to adapt. I have spent hours troubleshooting, only to find out it's because the documentation is referring to the same thing by two different names. It's 2024 also, I shouldn't be seeing print statements without parentheses.
Each thing seems to have a bunch of clicks to setup that startup LLM providers don't hassle people with. They're more likely to just let you sign in with some generic third party oAuth, slap on Stripe billing, let you generate keys, show you some usage stats, getting started docs, with example queries and a prompt playground etc.
What about the Vertex models though? Are they all actually available via Google AI Studio?
Billing for the Gemini models (on Vertex AI, the Generative Language AI variant still charges by tokens) I would argue is simpler than every other provider, simply because you're charged by characters/image/video-second/audio-second and don't need to run a tokenizer (if it's even available cough Claude 3 and Gemini) and having to figure out what the chat template is to calculate the token cost per message [2] or figure out how to calculate tokens for an image [3] to get cost estimates before actually submitting the request and getting usage info back.
[1]: https://cloud.google.com/vertex-ai/generative-ai/docs/multim...
[2]: https://platform.openai.com/docs/guides/text-generation/mana...
[3]: https://platform.openai.com/docs/guides/vision/calculating-c...
I'm also not sure if I understand your problem with pricing? Depending on what you do with it, it's not just an LLM. It actually started before llms.
Pricing for image classification and other features are completely different products like an LLM.
I use gcp professional every day and always found it quite intuitive.
Did plenty of image classification with vertex ai too
It's a shitty solution to a stupid problem ;)
But I did mention that vertex AI is more than just hosting llms though
Is this a contradiction or am I misunderstanding something?
Btw overall very impressive work great job.
However, I wouldn't draw conclusions about different model families, like Llama and Gemma, based on their token count alone. There are many other variables at play - the quality of those tokens, number of epochs, model architecture, hyperparameters, distillation, etc. that will have an influence on training efficiency.
The Google API models support 1M+ tokens, but these are just 8K. Is there a fundamental architecture difference, training set, something else?
(table 13 on page 7) vs https://arxiv.org/pdf/2404.14219 (page 6, quite better in general)
The report on knowledge distillation training is interesting, though.
The 2.6B would get stomped by Phi-3, so there's no comparison.
Fair enough. 2.6B vs. 3.8B is a fairly substantial size difference thats hard to intuit when its 2.6 vs 3.8 versus 2,600,000,000 and 3,800,000,000.
But then we get what I'm going to "parameter creep": Mistral 7B vs. Llama 8B vs. Gemma 9B. I worried after Llama 3 went 8B that we'd start seeing games with parameters, but, thought I was being silly.
The implication in my post is "if the reason was size, it's invalidated later"
https://aistudio.google.com/app/prompts/new_chat?model=gemma...
So far it seems pretty strong for its size.
That said, do you have a reason for keeping msty closed source rather than open? I read your FAQ for "why should I trust msty" and it feels lacking.
> We are a small team of developers who are passionate about AI and privacy. We have worked on projects before that have been used by thousands of people such as this (I've never heard of Cleavr). There are real faces (real faces = Twitter account link?) behind the product. And come chat with us on our Discord server to know us better.
This is much, much better than having no attribution, but it's miles away from being able to verify trust by reading the code. Would love to hear what your reasons against this are.
Still thinking about trying it out, anyway...
Trying to save Anthropic API key on Arch Linux doesn't do anything and there's a message "If you're experiencing problems saving API keys especially on Linux, contact Discord", if it's so common problem maybe you should have a link with possible fixes? Adding another Discord server and searching for answers for a question that clearly has been asked often enough feels like quite a hurdle for testing it out.
But I'm not seeing Gemma 2 or Claude 3.5 Sonnet even though it's announced on your landing page.
It does seem to be true that clean data works better than low quality data.
Model collapse itself is(was?) a fairly serious research topic: https://arxiv.org/abs/2305.17493
We've by now reached a "probably not inevitable" - https://arxiv.org/abs/2404.01413 argues there's a finite upper bound to error - but I'd also point out that that paper assumes training data cardinality increases with the number of training generations and is strictly accumulative.
To a first order, that means you better have a pre-2022 dataset to get started, and have archived it well.
but it's probably fair to say current SOTA is still more or less "it's neither impossible nor inevitable".
> To a first order, that means you better have a pre-2022 dataset to get started, and have archived it well.
I think that will always be available, or at least, a dataset with the distribution you want will be available.
>I think that [clean pre-2022 data set] will always be available
Good luck obtaining one.
Do I have to manually sanitize the input before I give it to the model?
See for example, the llama3 tokenizer has options to control special token tokenization:
Tokenization method with args to control special token handling: https://github.com/meta-llama/llama3/blob/bf8d18cd087a4a0b3f...
And you can see how it is used combined with special tokens and user input here: https://github.com/meta-llama/llama3/blob/bf8d18cd087a4a0b3f...
If you don't have control of the tokenizer, I guess it needs to be sanitized in the input like you say.
Hmmm. I'd love to know what qualifies as "unsafe".
I've seen documentaries and science shows on cable TV that demonstrate basic facts like this, or how the IRA produced IEDs, or how molotov cocktails were made in the spanish civil war.
The information is beyond easy to access, and has been for decades.
We saw all the bad press companies have got in recent years for all kinds of unintended AI outputs.
Porque no los dos?
Small models are never going to be generalists, so having several small models allows you to pick the one that best fits your needs.
I've found Gemini to be better at some use-cases, and GPT-4 better at others for my specific taste and use-case. You can kind of go by the benchmark scores to have an idea if it's good at logic, creativity, etc.
I imagine Gemma 2 is a better general-purpose assistant for most people, whereas Phi 3 is a solid small LLM (SLM?) for more specific use-cases like summarization, RAG, learning about math and stuff.
Gemma's performance if anything seems understated on benchmarks: the 27b is currently ahead of Llama3-70b on the Chatbot Arena leaderboard.
Guess I'll just fully test it for my own tasks to know for sure
And when we continue fine-tune.how much and what type of data we learn it on, I'm pretty sure for a smart agent who is not a knowledgeable expert but primarily a agent (understand what and how) this will get smaller and easier to run everywhere.
> User turn: user
> Model turn: model
> Start of conversation turn: <start_of_turn>
> End of conversation turn: <end_of_turn>
> Beginning of sequence: <bos>
> End of sequence: <eos>
You know I keep wondering why <bos> and <eos> tokens are even a thing in general. No model is tuned to keep generating multiple turns after its <end_of_turn> equivalent is sent, and what's the point of <bos> when you're parsing the entire context anyway. If it's an attempt to ignore text before it... then why is that text there? Just remove it from context, you're throwing away compute.
If you have a bunch of small prompts/answers, you can fit them into bigger batches if you use start/stop tokens.
To compensate for that, you can pack multiple examples in the same sequence. This is there EOS and BOS come in, as they indicate to the model that the two parts of the sequence are not related.
Why does using distillation from a larger model simulate training with more tokens?
Essentially instead of tokens that are "already there" in text, the distillation allows us to simulate training data from a larger model
What are the theories as to why this works better than training on a larger quantity of non-simulated tokens?
Is it because the gradient from the non-simulated tokens is too noisy for a small model to model correctly?
```python def solve_quadratic_equation(a, b, c): """Solves a quadratic equation of the form ax^2 + bx + c = 0."""
discriminant = (b ** 2) - (4 * a * a)
if discriminant >= 0:
root = (-b + math.sqrt(b ** 2 - 4 * a * a ** b**
0.5 #
1.
# Return None if the quadratic equation has no real roots
if (b ** 2) < (4 * c):
return None
# Calculate the roots using the quadratic formula
b = -b
b
# a, b): Solve for the discriminant.
# Handle the case of a complex discriminant
# Print the solution to the equation
if (b * 2)print("The quadratic equation is: " + a * x* 2 + b "x" + c) ```
Are the weights released yet?
for basic llm tasks that most people would use on their daily lives (simple rag on your own data), it did the job for the most part (unless you need a lot of context maybe).
on paper the newer one shows significant improvement with slightly larger size, but i hope HumanEval regression is not going to matter for most people.
Benchmark | Gemma 2 (9B) | Phi-3 Small (7B)
-----------------------------|----------------|-------------------
MMLU (5-Shot) | 63.6 | 75.7
HellaSwag (5-Shot) | 49.8 | 77.0
ANLI (7-Shot) | 48.7 | 58.1
GSM-8K (8-Shot; CoT) | 59.8 | 89.6
MedQA (2-Shot) | 49.6 | 65.4
AGIEval (0-Shot) | 42.1 | 45.1
TriviaQA (5-Shot) | 72.3 | 58.1
Arc-C (10-Shot) | 78.3 | 90.7
Arc-E (10-Shot) | 91.4 | 97.0
PIQA (5-Shot) | 78.1 | 86.9
SociQA (5-Shot) | 65.5 | 79.2
BigBench-Hard (3-Shot; CoT) | 59.6 | 79.1
WinoGrande (5-Shot) | 55.6 | 81.5
OpenBookQA (10-Shot) | 78.6 | 88.0
BoolQ (2-Shot) | 66.0 | 84.8
CommonSenseQA (10-Shot) | 76.2 | 80.0
TruthfulQA (10-Shot; MC2) | 52.1 | 70.2
HumanEval (0-Shot) | 34.1 | 61.0
MBPP (3-Shot) | 51.5 | 71.7I think Google just lacks the vision to understand what makes a good LLM. Theoretical contributions by research teams are valuable, but the real-world is built around engineering ideas that may lack the "purity" and elegance of theory but damn it they work.
[1]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
I gave up on Chrome a decade ago, going back to Firefox. I don't use Google for search anymore, I do use Gmail but I also got Protonmail so could easily migrate the Gmail traffic there.
A lot of non-techies I know have complained for some time how Google search sucks, and while a lot use Chrome it seems to be mainly inertia.
Not saying Google is dying, but it seems vulnerable for disruption.
I wouldn't be surprised if most of it was already tried in some form.
Claude's reign has begun, and I'd say it has a solid enough lead for at least another two weeks of dominance before it's dethroned.
This is an incredible statement to make about a field that no one was talking about 24 months ago, a family of SOTA models that didn't exist until 8 months ago, and a family of small local models that didn't exist 6 months ago. But sure, give up hope after the first generation of a model family doesn't impress you.
People seem to forget how incredibly early we are in this whole thing. The fact that so much progress has been made in such a short amount of time should make everyone super excited!
That's not to say this is an insignificant contribution. New models are great, especially when released for free, and it's important for big firms to keep the ball rolling for tech to progress. Though there is also legitimate concern that all LLMs aren't improving as fast as they used to improve, and we may have hit the proverbial bathtub curve of AI progress.