HNHacker News
TopNewBestAskShowJobs

kathleenfromgdm

97 karma · joined February 21, 2024

submissionscomments
kathleenfromgdm··on Gemma: New Open Models
The context length for these models is 8192 tokens.
kathleenfromgdm··on Gemma: New Open Models
Good catch - just corrected. Thanks!
kathleenfromgdm··on Gemma: New Open Models
The context length for these models is 8192 tokens.
kathleenfromgdm··on Gemma: New Open Models
Great question - we compare to the Mistral 7B 0.1 pretrained models (since there were no pretrained checkpoint updates in 0.2) and the Mistral 7B 0.2 instruction-tuned models in the technical report here: https://goo.gle/GemmaReport
kathleenfromgdm··on Gemma: New Open Models
We release our non-aligned models (marked as pretrained or PT models across platforms) alongside our fine-tuned checkpoints; for example, here is our pretrained 7B checkpoint for download: https://www.kaggle.com/models/google/gemma/frameworks/keras/...
kathleenfromgdm··on Gemma: New Open Models
Corrected - thanks :)
kathleenfromgdm··on Gemma: New Open Models
Yes, you can get started downloading the model and running inference on Kaggle: https://www.kaggle.com/models/google/gemma ; for a full list of ways to interact with the model, you can check out https://ai.google.dev/gemma.
kathleenfromgdm··on Gemma: New Open Models
Thank you! You can get started downloading the model and running inference on Kaggle: https://www.kaggle.com/models/google/gemma ; for a full list of ways to interact with the model, you can check out https://ai.google.dev/gemma.
kathleenfromgdm··on Gemma: New Open Models
We've documented the architecture (including key differences) in our technical report here (https://goo.gle/GemmaReport), and you can see the architecture implementation in our Git Repo (https://github.com/google-deepmind/gemma).