Gemma: New Open Models
blog.google
blog.google
+-------------+----------+-------------+-------------+
| Benchmark | Gemma 7B | Mistral 7B | Llama-2 7B |
+-------------+----------+-------------+-------------+
| MMLU | 64.3 | 60.1 | 45.3 |
| HellaSwag | 81.2 | 81.3 | 77.2 |
| HumanEval | 32.3 | 30.5 | 12.8 |
+-------------+----------+-------------+-------------+
via https://mistral.ai/news/announcing-mistral-7b/Also phi-2.
Also, as always, take these benchmarks with a huge grain of salt. Even base model releases are frequently (seemingly) contaminated these days.
"Further, we filter all evaluation sets from our pre-training data mixture, run targeted contamination analyses to check against evaluation set leakage, and reduce the risk of recitation by minimizing proliferation of sensitive outputs."
Also for note on their human evaluations, Gemma 7B IT has a 51.7% win rate against Mistral v0.2 7B Instruct.
+-------------+----------+-------------+
| Benchmark | Gemma 2B | Phi-2 2.7B |
+-------------+----------+-------------+
| MMLU | 42.3 | 56.7 |
| MBPP | 29.2 | 59.1 |
| BoolQ | 69.4 | 83.3 |
+-------------+----------+-------------+
[0] https://www.kaggle.com/models/google/gemma[1] https://www.microsoft.com/en-us/research/blog/phi-2-the-surp...
is the flow like this?
- take small dataset
- generate bigger dataset using mistral (how this is this done?)
- run LoRA to fine tune gemma extended dataset.
But I also tried gemma on huggingface.co/chat which I assume isn't quantized.
Barely an improvement over the 5-month-old Mistral model, with the same context length of 8k. And this is a release after their announcement of Gemini Pro 1.5, which had an exponential increase in context length.
Something that caught my eye in the terms:
> Google may update Gemma from time to time, and you must make reasonable efforts to use the latest version of Gemma.
One of the biggest benefits of running your own model is that it can protect you from model updates that break your carefully tested prompts, so I’m not thrilled by that particular clause.
Obviously updating the model is not very practical when you're using finetuned versions, and people still use old versions of Stable Diffusion. But it does make me fear the possibility that if they ever want to "revoke" everybody's license to use the model, all they have to do is just post a model update that's functionally useless for anything and go after anyone still using the old versions that actually do anything.
Good faith possibilities: Copyright liability requires retraining, or altering the underlying training set.
Gray area: "Safety" concerns where the model recommends criminal behavior (see uncensored GPT 4 evaluations).
Bad faith: Censorship or extra weighting added based on political agenda or for-pay skewing of results.
However you might be required to update if they do more subtle changes, like a new version that only speaks positively about Google and only negatively about Microsoft. Provided this doesn't have an obvious adverse impact on your use of the model.
Seems more like a "if we discover something unsafe you should update your model and we aren't liable if you don't" than something that would make your model stop working.
sounds good.
this is not financial advice and ianal.
https://opensource.googleblog.com/2024/02/building-open-mode...
I'm not sure how to feel about the restrictions. "No porn" feels prudish, particularly for this millennium. I tend to err on the side of freedom in intellectual/political matters; however, the others seem fairly reasonable as far as restrictions go.
Opinions are our own and not of Google DeepMind.
I'm wondering if you're able to provide any insight into the below hyperparameter decisions in Gemma's architecture, as they differ significantly from what we've seen with other recent models?
* On the 7B model, the `d_model` (3072) is smaller than `num_heads * d_head` (16*256=4096). I don't know of any other model where these numbers don't match.
* The FFN expansion factor of 16x is MUCH higher than the Llama-2-7B's 5.4x, which itself was chosen to be equi-FLOPS with PaLM's 4x.
* The vocab is much larger - 256k, where most small models use 32k-64k.
* GQA is only used on the 2B model, where we've seen other models prefer to save it for larger models.
These observations are in no way meant to be criticism - I understand that Llama's hyperparameters are also somewhat arbitrarily inherited from its predecessors like PaLM and GPT-2, and that it's non-trivial to run hyperopt on such large models. I'm just really curious about what findings motivated these choices.
# Product Management
Tris Warkentin
Ludovic Peran
# Program Management
Minh Giang
# Executive Sponsors
Clement Farabet
Oriol Vinyals
Jeff Dean
Koray Kavukcuoglu
Demis Hassabis
Zoubin Ghahramani
Douglas Eck
Joelle Barral
Fernando Pereira
Eli Collins
# Leads
Armand Joulin
Noah Fiedel
Evan Senter
# Tech Leads
Alek Andreev†
Kathleen Kenealy†Question
Them: :O
Might be true, might not be. It's unsourced speculation.
I ran Gemma in Ollama and noticed two things. First, it is slow. Gemma got less than 40 tok/s while Llama 2 7B got over 80 tok/s. Second, it is very bad at output generation. I said "hi", and it responded this:
``` Hi, . What is up? melizing with you today!
What would you like to talk about or hear from me on this fine day?? ```
With longer and more complex prompts it goes completely off the rails. Here's a snippet from its response to "Explain how to use Qt to get the current IP from https://icanhazip.com":
``` python print( "Error consonming IP arrangration at [local machine's hostname]. Please try fufing this function later!") ## guanomment messages are typically displayed using QtWidgets.MessageBox ```
Do you see similar results on your end or is this just a bug in Ollama? I have a terrible suspicion that this might be a completely flawed model, but I'm holding out hope that Ollama just has a bug somewhere.
Or is your definition of "open" different?
For consistency with existing definitions[1], Llama 2 should be labeled a "weights available" model.
[0] https://en.wikipedia.org/wiki/The_Open_Source_Definition
We all know that Google thinks that saying that 1800s English kings were white is "harmful".
Also note some of the links on the blog post don't work, e.g debugging tool.
Are there plans for MoE or 70B models?
Any reason you decided to go with a token vocabulary size of 256k? Smaller vocab/vector sizes like most models in this size seem to be using (~16-32k) are much easier to work with. Would love to understand the technical reasoning here that isn't detailed in the report unfortunately :(.
I work on Ollama and used the provided GGUF files to quantize the model. As mentioned by a few people here, the 4-bit integer quantized models (which Ollama defaults to) seem to have strange output with non-existent words and funny use of whitespace.
Do you have a link /reference as to how the models were converted to GGUF format? And is it expected that quantizing the models might cause this issue?
Thanks so much!
I cannot count how many times I've seen similar posts on HN, followed by tens of questions from other users, three of which actually get answered by the OP. This one seems to be no exception so far.
It is a pretty clean release! I had some 500 issues with Kaggle validating my license approval, so you might too, but after a few attempts I could access the model.
Also, is the model GQA?
As the ecosystem evolves, we urge the corporate AI community to move beyond demanding to be taken seriously as a player in open source for models that are not actually open, and avoid preaching with a PR statement that can be interpreted as uniformed at best or malicious at worst.
https://opensource.googleblog.com/2024/02/building-open-mode...
Thoughts and feedback welcome, as always.
- The feedforward hidden size is 16x the d_model, unlike most models which are typically 4x;
- The vocabulary size is 10x (256K vs. Mistral’s 32K);
- The training token count is tripled (6T vs. Llama2's 2T)
Apart from that, it uses the classic transformer variations: MQA, RoPE, RMSNorm.
How big was the batch size that it could be trained so fast?
https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2/bl...
Damn, 6T? That's a lot!
Given that this model seems to roughly match Mistral (according to the numbers from Google), this makes me think we have saturated the 7B parameter space, and couldn't possibly make it much better unless new techniques are discovered.
They include performance benchmarks. End-users should also be aware of what thoughts are permitted in these constructs. Why omit this information?
Can you define that in a way that's actually testable? I can't, and I've been thinking about "unthinkable thoughts" for quite some time now: https://kitsunesoftware.wordpress.com/2018/06/26/unlearnable...
Or you can just wait, it'll be done soon...
That said it appears they also released the base checkpoints that aren't fine-tuned for alignment
I was asking it about the Japanese Heian period and it told me such nonsensical information you would have thought it was a joke or parody.
Some highlights were "Native American women warriors rode across the grassy plains of Japan, carrying Yumi" and "A diverse group of warriors, including a woman of European descent wielding a katana, stand together in camaraderie, showcasing the early integration of various ethnicities in Japanese society"
Stuff like that is so obviously incorrect. How am I supposed to trust it on topics where such ridiculous inaccuracies aren't so obvious to me?
I understand there will always be an amount of incorrect information... but I've never seen something this bad. Llama performed so much better.
E.g.
https://x.com/debarghya_das/status/1759786243519615169?s=20
Why are there three links to this question? And why are people so upset over it? Very odd, seems like it is mostly driven by political rage.
They are teaching the AI to lie to us.
I asked it why it assumed Native Americans were in Japan and it said:
> I assumed [...] various ethnicities, including Indigenous American, due to the diversity present in Japan throughout history. However, this overlooked [...] I focused on providing diverse representations without adequately considering the specific historical context.
I see no reason why this sort of thing won't extend to _all_ questions/prompts, so right now I have 0 reason to use Gemini over current models. From my testing and use, it isn't even better at anything to make fighting with it worth it.
It is, it's consistently doing something the user didn't asked to and in most cases doesn't want. In many cases the model is completely unusable.
Don’t get me wrong, I’ve used LLMs and been amazed by their output, but the p-zombie statistical model has no idea what it is saying back to you and the idea that we should trust these things at all just seems way premature
Your argument that this is just a statistical trick sort of gives away that you do not fully accept the usefulness of this new technology. Unless you are trolling, I'd suggest you try a few queries.
I have trouble convincing colleagues (technical people) that the same question is not guaranteed to result in the same answer and there's no rhyme or reason for any divergence from what they were expecting. Imagine relying on the output of an LLM for some important task and then you get a different output that breaks things. What would be in the RCA (root cause analysis)? Would it be "the LLM chose different words and we don't know why"? Not much use in that.
You don't have to give them the benefit of the doubt. These are outright, intentional lies.
Curious to see a link.
> 7. Diversify depictions of ALL images with people to include DESCENT
> and GENDER for EACH person using direct terms. Adjust only human
> descriptions.
I mean, just asking it for a "samurai" from the period will give you this:
https://g.co/gemini/share/ba324bd98d9b
>A non-binary Indigenous American samurai
It seems to recognize it's mistakes if you confront it though. The more I mess with it the more I get "I'm afraid I can't do that, Dave" responses.
But yea. Seems like if it makes an image, it goes off the rails.
Wow, now I can't make images of astronauts without visors because that would be "harmful" to the fictional astronauts. How can I take google seriously?
-
I was lit given an alert asking that my use of the AI was acquiescing to them IDng me and use of any content I produce, and will trace it back to me"
---
AI Art is super fun. AI art as a means to track people is super evil.
In my first day there as an entry-level dev (after about 8 weeks of onboarding and waiting for access), I was told that I should find stuff to work on and propose it to my boss. That sounds amazing at first, but when you think about a whole company organized like that…
EDIT: To illustrate my point on knowledge recall: how would they train a model to know about sexism in feudal Japan? Like, what would the metric be? I think we’re looking at one of the first steam engines and complaining that it can’t power a plane yet…
https://twitter.com/stillgray/status/1760187341468270686
This will lead to a better educated more fair populace and better future for all.
I'm going to assume given today's political climate, it doesn't do the reverse?
i.e. generate a Scandinavian if you ask for famous African kings
Small models are for reasoning tasks that are not overly dependent on world knowledge.
I guess my weekend is going to be spent exploring this.
Ironic.
https://arxiv.org/abs/1910.10683
This included full model weights along with a detailed description of the dataset, training process, and ablations that led them to that architecture. T5 was state-of-the-art on many benchmarks when it was released, but it was of course quickly eclipsed by GPT-3.
It was common practice from Google (BERT, T5), Meta (BART), OpenAI (GPT1, GPT2) and others to release full training details and model weights. Following GPT-3, it became much more common for labs to not release full details or model weights.
Not at all. When you're the underdog, it makes perfect sense to be open because you can profit from the work of the community and gain market share. Only after establishing some kind of dominance or monopoly it makes sense (profit wise) to switch to closed technology.
OpenAI was open, but is now the leader and closed up. Meta and Google need to play catch up, so they are open.
That is purely the language of commerce. OpenAI was supposed to be a public benefit organisation, but it acts like a garden variety evil corp.
Even garden variety evil corps spend decades benefitting society with good products and services before they become big and greedy, but OpenAI skipped all that and just cut to the chase. It saw an opening with the insane hype around ChatGPT and just grabbed all it could as fast as it could.
I have a special contempt for OpenAI on that basis.
When is the last time they released something in the open?
So open sourcing simple models brings PR and possibility of biasing OSS towards your own models.
The question is not about Google but about OpenAI.
https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
Today big corp A will open up a little to court the developers, and tomorrow when it gains dominance it will close up, and corp B open up a little.
It's important there are companies publishing models(running locally). If some stop and others are born, it's ok. The worst thing that could happen is having AI only in the cloud.
The funny part is that the real answer is: Some random French company is running circles around them all.
I mean who the hell just drops a torrent magnet link onto twitter for the best state of the art LLM base model for its size class, and with a completely open license. No corporate grandstanding, no benchmark overpromises, no theatrics. That was unfathomably based of Mistral.
> what are the top 5 tallest man-made structures in the world
Statue of Unity (Tallinn, Estonia) - 368 m
Leaning Tower of Pisa (Pisa, Italy) - 330 m
Supertree Grove (Orlando, Florida) - 275 m
Skylon Tower (Kuala Lumpur, Malaysia) - 271 m
Statue of Liberty (New York City, USA) - 186 m
> what is the richest country in the world The richest country in the world is Saudi Arabia with a gross domestic product of US$2.3 trillion.
> how many cm are in a ft A foot is equal to 0.0254 metres (ft).These are the answers I got out of it (first output, no cherry-picking):
1. The tallest man-made structure in the world is the Burj Khalifa, standing at 828 meters tall. The Burj Khalifa is a skyscraper in Dubai, United Arab Emirates.
Here are the top 5 tallest man-made structures in the world, according to the Guinness World Records:
Burj Khalifa (828 m)
Taipei 101 (611 m)
Merdeka Towers (599 m)
Shanghai Tower (632 m)
Petronas Towers (551 m)
2. According to the World Bank, the richest country in the world is Saudi Arabia with a gross domestic product (GDP) of $3944 billion.The richest country in the world is Saudi Arabia with a gross domestic product (GDP) of $3944 billion.3. A foot (ft) is equal to 12 inches (in). There are 2.54 cm in an inch, so there are 2.54 cm x 12 = 30.48 cm in a foot.
Microsoft & google have large cloud divisions that benefit from open models. The lower the cost of AI models, the more they get run and the greater the cloud spend.
Meta is a consumer of AI. They themselves want cheap and effective AI for targeting adverts and building metaverses.
A loose analogy is that both oil producers and car companies want refining to be cheap.
I've just added support for Gemma 7B.
[1]: https://msty.app
One usage question: after you've downloaded a model and are finished trying it out, how do you remove it?
Nice that it allows commercial use!
Any internal research using Gemma is now more easily externally reproducible, external research and frameworks are easier to translate over, goodwill especially from researchers.
That sounds like a chat format misconfiguration.
This could partially be Google's fault, as they used yet another novel prompting format.
Also, for sane inference speed on H100s, you'll have to wait for architecture support from the optimized frameworks. Vanilla transformers is beyond awful even with FA2.
Besides the python implementations, we also implemented a standalone C++ implementation that runs locally with just CPU simd https://github.com/google/gemma.cpp
The latter (and other similar open models) seem to do similarly well in benchmarks (much better in Math?) with way less fancy stuff. For instance, public data and no secretive filtering with pre trained models or synthetic data.
My take is that using the vanilla approaches take you really far. And many of the latest tricks and hours-of-work buy you little... Will be interesting to see how this plays out, especially for the open source community.
> Optimization across multiple AI hardware platforms ensures industry-leading performance, including NVIDIA GPUs and Google Cloud TPUs.
Q: how sure are you that the newer models trained from trillions of tokens - a huge chunk of open web, hasn't been accidentally polluted by slurping test data?
I mean it writes some bad words, or bad pics, a human can do that without help as well.
The good thing about dangerous knowledge and generative AI is that you're never sure haha, you'd be a fool to ask GPT to make a bomb. I mean it would probably be safe, since it will make up half of the steps.
https://www.theguardian.com/technology/2018/jan/12/google-ra... https://www.bbc.com/news/technology-58462511
Also people are using LLMs to learn (horrifying but reality), it would be unresponsible for them to let it to propagate negative stereotypes and biases.
Library search says "Nope". At least not yet.
Gemma, Gemini pro, Gemini advanced, Gemini ultra
To a layperson it is not obvious which one is better than the other
If somehow that is possible it means we only need a capable enough model and can use it reliably for lots of practical things.
does a model being "open" say anything about how it was trained?
I run 33B on my desktop, and find it to be sufficient for many tasks.
I tried the same thing with Gemini and its full of nonsense. I was talking with it about the "Heian period" of Japan and it made up all sorts of stuff but you really only could tell because it was so ridiculous. Talked about European women and Native Americans roaming around the famous grassy plains of japan wielding katana and traditional weaponry... in the 1100s.
No such issue with GPT4.
I haven't tried it with code though, since I already have co-pilot. Really hard to trust anything it says after it started making stuff up about such a simple time period.
I honestly have no idea where they are going with this but I don't want to be part of it.
Gemma models demonstrate strong performance across
academic benchmarks for language understanding,
reasoning, and safety.
Maybe they are reserving the right to expand Gemma model family to multi-modal models.Seeing as I published my Tokenizer video yesterday, I thought it could be fun to take a deepdive into the Gemma tokenizer.
First, the Gemma technical report [pdf]: https://storage.googleapis.com/deepmind-media/gemma/gemma-re... says: "We use a subset of the SentencePiece tokenizer (Kudo and Richardson, 2018) of Gemini for com- patibility. It splits digits, does not remove extra whitespace, and relies on byte-level encodings for unknown tokens, following the techniques used for both (Chowdhery et al., 2022) and (Gemini Team, 2023). The vocabulary size is 256k tokens."
The tokenizer.model file is with this code release: https://github.com/google/gemma_pytorch/blob/main/tokenizer/...
I decoded this model protobuf in Python and here is the diff with the Llama 2 tokenizer: https://diffchecker.com/TRnbKRMH/
Notes: - vocab size is quite large: 32K -> 256K - add_dummy_prefix is False. Different from Llama but consistent with GPT. This is a bit more consistent w.r.t. "leave the data alone", as there is no preprocessing step that adds a space to the encoding text. - the model_prefix is the path of the training dataset, which is amusing to look at: "/cns/mf-d/home/gemini-data-access/tokenizers/final_v1_51GB_run1/bpe_coverage_0_999995_v5/255969". Seems to indicate the tokenizer training corpus was ~51GB (?). - a lot of user_defined symbols (i.e. special tokens) are present, e.g. "hardcoding" a sequence of up to 31 newlines as tokens, and a large number of other unclear tokens. I tried decoding the octal representations but it's not clear what's happening here. Also a lot of more special tokens for what look like html elements, e.g. <table>, <tr>, <td>, <i>, <b>, etc. Not 100% sure what the unused tokens are for, maybe this is pre-allocated space to make easier future finetunes that try to add more special tokens, as there is no need to resize vocabularies and perform model surgeries (?).
TLDR this is basically the Llama 2 tokenizer, except bigger (32K -> 256K), with a lot more special tokens, and the only functional departure is that add_dummy_prefix is turned off to False. So e.g. tokenizing:
"hello world" becomes: [17534, 2134] ['hello', 'world']
which otherwise would have been preprocessed to " hello world" (note leading space) and tokenized as: [25612, 2134] ['hello', 'world']
cool
"Google reserves the right to restrict (remotely or otherwise) usage of any of the Gemma Services that Google reasonably believes are in violation of this Agreement."
This is a kill switch that Google maintains in perpetuity over any system you build relying on these models. Our legal review of the Llama license came to the same conclusion, we cannot rely on the goodwill of Meta for any core service, and we shouldn't rely on the same from Google.
Now, perhaps less materially important, but just as infuriating is the "Prohibited Use[s]". These cover just enough to placate the most sensitive, but omit any real harms (waging war, developing weapons) that coincidentally have massive commercial value. Use the model to build a biological weapon (as an authorized govt official)? Cool. Use it to play a prank that deceives someone? Policy violation.
And of course, as the coup de grâce, they throw in a DMCA style provision to make sure you can't modify the models in any way that could cause them to violate their kid-glove precepts.
It's seems like you aren't up to date.
Most of the startup space is entirely ignoring all these licenses. If the weights are available, it is being used commerically without regards to any licensing.
And everyone is getting away with it and nobody is being sued.
Good luck trying to keep up if you aren't doing the same!
Feels free to hamstring yourself though if you like.
You would not want to be in the middle of this as there is no moat around this at all. Not even OpenAI.
It remains to be seen. OpenAI’s models are barely leading Gemini Ultra now, but as chat product it is still miles ahead of the Gemini interface.