Gemma.cpp: lightweight, standalone C++ inference engine for Gemma models
github.com
github.com
To get a few common questions out of the way:
- This is separate / independent of llama.cpp / ggml. I'm a big fan of that project and it was an inspiration (we say as much in the README). I've been a big advocate of gguf + llama.cpp support for gemma and am happy for people to use that.
- how is it different than inference runtime X? gemma.cpp is a direct implementation of gemma, in its current form it's aimed at experimentation + research and portability + easy modifiable rather than a general purpose deployment framework.
- this initial implementation is cpu simd centric. we're exploring options for portable gpu support but the cool thing is it will build and run on a lot of environments you might not expect an llm to run, so long as you have the memory to load the model.
- I'll let other colleagues answer questions about the Gemma model itself, this is a C++ implementation of the model, but relatively independent of the model training process.
- Although this is from Google, we're a very small team that wanted such a codebase to exist. We have lots of plans to use it ourselves and we hope other people like it and find it useful.
- I wrote a twitter thread on this project here: https://twitter.com/austinvhuang/status/1760375890448429459
In its current form, I think of gemma.cpp is more of a direct model implementation (somewhere between the minimalism of llama2.c and the generality of ggml).
I tend to think of 3 modes of usage:
- hacking on inference internals - there's very little indirection, no IRs, the model is just code, so if you want to add support for your own runtime support for sparsity/quantization/model compression/etc. and demo it working with gemma, there's minimal barriers to do so
- implementing experimental frontends - i'll add some examples of this in the very near future. but you're free to get pretty creative with terminal UIs, code that interact with model internals like the KV cache, accepting/rejecting tokens etc.
- interacting with the model locally with a small program - of course there's other options for this but hopefully this is one way to play with gemma w/ minimal fuss.
That sounds interesting
ps i'm a fan of cosmopolitan as well.
This is really cool, Austin. Kudos to your team!
Everyone working on this self-selected into contributing, so I think of it less as my team than ... a team?
Specifically want to call out: Jan Wassenberg (author of https://github.com/google/highway) and I started gemma.cpp as a small project just a few months ago + Phil Culliton, Dan Zheng, and Paul Chang + of course the GDM Gemma team.
Austin and Jan are truly amazing. The optimization work is genuinely outstanding; I get incredible CPU performance on Gemma.cpp for inference. Thanks for all of the awesomeness, Austin =)
python wrapper - if you want to run the model in python I feel like there's already a lot of more mature options available (see the model variations at https://www.kaggle.com/models/google/gemma) , but if people really want this and have something they want to do with a python wrapper that can't be done with existing options let me know. (similar thoughts wrt to API servers).
I would think that having an OpenAI standard compatible API would be a higher priority over a python wrapper, since then it can act as a drop in replacement for most any backend.
- Somewhere the README, consider adding the need for a `-DWEIGHT_TYPE=hwy::bfloat16_t` flag for non-sfp. Maybe around step 3.
- The README should explicitly say somehere that there's no GPU support (at the moment)
- "Failed to read cache gating_ein_0 (error 294)" is pretty obscure. I think even "(error at line number 294)" would be a big improvement when it fails to FindKey.
- There's something odd about the 2b vs 7b model. The 2b will claim its trained by Google but the 7b won't. Were these trained on the same data?
- Are the .sbs weights the same weights as the GGUF? I'm getting different answers compared to llama.cpp. Do you know of a good way to compare the two? Any way to make both deterministic? Or even dump probability distributions on the first (or any) token to compare?
The weights should be the same across formats, but it's easy for differences to arise due to quantization and/or subtle implementation differences. Minor implementation differences has been a pain point in the ML ecosystem for a while (w/ IRs, onnx, python vs. runtime, etc.), but hopefully the differences aren't too significant (if they are, it's a bug in one of the implementations).
There were quantization fixes like https://twitter.com/ggerganov/status/1760418864418934922 and other patches happening, but it may take a few days for patches to work their way through the ecosystem.
I'm using the 32-bit GGUF model from the Google repo, not a different quantized model, so I could have one less source of error. It's hard to tell with LLMs if its a bug. It just gives slightly stranger answers sometimes, but it's not completely gibberish. or incoherent sentences or have extra punctuations like with some other LLM bugs I've seen.
Still, I'll wait a few days to build llama.cpp again to see if there are any changes.
1) Reference implementations in JAX, PyTorch, TF with Keras 3, MaxText/JAX, more...
2) Full integration at launch with HF including Transformers + optimization therein
3) TensorRT-LLM and full NVIDIA opt across the stack in partnership with that team (mentioned on the NVIDIA earnings call by Jensen, even)
4) More developer surfaces than you can shake a stick at: Kaggle, Colab, Gemma.cpp, GGUF
5) Comms landing with full coordination from Sundar + Demis + Jeff Dean, not to mention positive articles in NYT, Verge, Fortune, etc.
6) Full Google Cloud launches across several major products, including Vertex and GKE
7) Launched globally and with a permissive set of terms that enable developers to do awesome stuff
Pulling that off without any major SNAFUs is a huge relief for the team. We're excited by the potential of using all of those surfaces and the launch momentum to build a lot more great things for you all =)
Now, I’m off playing with a new toy :)
Gemma support has been added to llama.cpp, and we're more than happy to see people use it there.
side note: imagine how gnarly those urls would be if HN used UUIDs instead of integers for IDs :-D
But Google is scarily capable on the LLM front and we shouldn't count them out. OpenAI might have the advantage of being quick to move, but when the juggernaut gets passed its resting inertia and starts to gain momentum it is going to leave an impression.
That became clear to me after watching the recent Jeff Dean video [1] which was posted a few days ago. The depth of institutional knowledge that is going to be unlocked inside Google is actually frightening for me to consider.
I hope the continued competition on the open source front, which we can really thank Facebook and Llama for, keeps these behemoths sharing. As OpenAI moves further from its original mission into capitalizing on its technological lead, we have to remember why the original vision they had is important.
So thank you, Google, for this.
1. https://www.youtube.com/watch?v=oSCRZkSQ1CE&ab_channel=RiceK...
Google has had years to get to this stage, and they've lost a lot of the talent that made their initial big splashes to OAI and competitors. Try finding someone on a sparse MoE paper from Google prior to 2022 who is still working there and not at OAI.
With respect, they can hardly even beat Mistral, resorting to rounding down a 7.8b model (w/o embeddings) to 7b.
Google has been the home of the talent for many years. They came on my radar in the late 00s when I used Peter Norvig's textbook in college, and they hired Ray Kurzweil in like 2012 or 2013 IIRC. They were hiring ML PhDs with talent for many years, and they pioneered most of the major innovations. They just got behind on productizing and shipping.
And of course writing that gives me a terrible realization: product placement in LLMs is going to be a very big thing in the near future.
https://youtu.be/-i9AGk3DJ90?t=616
In essence, Google already rules information retrieval. Their margins are insane. Switching to LLM based search cuts into their margins and increases their costs dramatically. Also, the advantage they've built over decades has been cut down.
All of this means there is potential for less profit and a shrinking valuation. A shrinking valuation means issues with employee retention and it could lead to long term stagnation.
Google has as much or more computing power than anyone. They're massively capitalized and have a market cap of almost $2T and colossal cashflow, and have the ability to throw enormous resources at the problem until they have a competitor. They have an enormous, benchmark-setting amount of data across their various projects to train on. That we're talking like they're some scrappy upstart is super weird.
>As OpenAI moves further from its original mission into capitalizing on its technological lead, we have to remember why the original vision they had is important.
I'm way more cynical about the open source models released by the megas, and OpenAI is probably the most honest about their intentions. Meta and Google are releasing these models arguably to kneecap any possible next OpenAI. They want to basically set the market value of anything below state of the art at $0.00, ensuring that there is no breathing room below the $2T cos. These models (Llama, Gemma, etc) are fun toys, but in the end they're completely uncompetitive and will yield zero "wins", so to speak.
Never thought about it that way, but it makes a lot of sense. It’s also true these models are not up to par with SOTA no matter what the benchmarks say
Google has the technical resources to become a major player here, maybe even the dominant player. But it won't happen under current management. I won't count out Google entirely, and there's still time for the company to be saved. It starts with new leadership.
Wait, it's not a major player in ML/AI?
You should see Google's turnover numbers from 4 years ago, much less now.
It's been years, it's broken internally, we see the results.
Here, we're in awe of 1KLOC of C++ code that runs inference on the CPU.
Nobody serious is running inference on CPU unless you're on the extreme cutting edge. (ex. I need to on Android and on the Chrome OS Linux VM, but I still use llama.cpp because it does support GPU everywhere else)
I'm not sure what else to say.
(n.b. i am a xoogler)
This code is also intended to facilitate research & experimentation, which may not fall under your definition of 'serious' :)
High turnover was industry-wide a few years back because pay went through the roof and job hopping was the best way to capture that.
I suspect it’s lower now, following mass layoffs and substantially fewer openings.
the core difference is the core model's performance, but the maximum potential is night and day imho. they had a good chance to do well in this front but did not end up making the most yet.
I wasn't familiar with the term, good article - https://masterofcode.com/blog/hallucinations-in-llms-what-yo...
I just got into hobby projects with diffusion a week ago and I'm seeing non-stop releases. It's hard to keep up. It's a firehose of information, acronyms, code etc.
It's been a great python refresher.
In fact it's probably better to dive deep into one hobby project like you're doing than constantly context switch with every little news item that comes up.
While working on gemma.cpp there were definitely a lot of "gee i wish i could clone myself and work on that other thing too".
My personal way of understanding it is this - the original sin of model weight format complexity is that NNs are both data and computation.
Representing the computation as data is the hard part and that's where the simplicity falls apart. Do you embed the compute graph? If so, what do you do about different frameworks supporting overlapping but distinct operations. Do you need the artifact to make training reproducible? Well that's an even more complex computation that you have to serialize as data. And so on..
GGML and GGUF are the same thing, GGUF is the new version that adds more data about the model so it's easy to support multiple architectures, and also includes prompt templates. These can run CPU only, be partially or fully offloaded to a GPU. With K quants, you can get anywhere from a 2 bit to an 8 bit GGUF.
GPTQ was the GPU-only optimized quantization method that was superseded by AWQ, which is roughly 2x faster and now by EXL2 which is even better. These are usually only 4 bit.
Safetensors and pytorch bin files are raw float16 model files, these are only really used for continued fine tuning.
That sounds very convenient. What software makes use of the built-in prompt template?
GGUF is just weights, safetensors the same thing. GGUF doesn't need a JSON decoder for the format while safetensors needs that.
I personally think having a JSON decoder is not a big deal and make the format more amendable, given GGUF evolves too.
I'm using LLMs locally almost always and eschewing API backed LLMs like chatgpt. So I'm not very familiar with plugins, and I'm assuming chatgpt plugs into a backend when it detects a math problem. So it isn't the LLM doing the math but to the user it appears to be.
Does anyone here know what LLM projects like llama.cpp or gemma.cpp support a plugin model?
I'm interested in adding to the dungeons and dragons system I built using llama.cpp. Because it doesn't do math well, the combat mode is terrible. But I was writing my own layer to break out when combat mode occurs, and I'm wondering if there is a better way with some kind of plugin approach.
https://chub.ai/characters/creamsan/team-neko-e4f1b2f8
This one says it uses javascript as well:
https://chub.ai/characters/creamsan/tessa-c4b917f9
Thise are the only two listed as SFW. There's some others if you hit the nsfw toggle and search for the scripted tag.I don't know if this is the right approach but you could also write a module for Sillytavern Extras.
I would like to inquire about the accessibility of the code and weights for the Gemma Model. Is this information publicly available?
I'm assuming you mean in other languages/implementations? (since the gemma.cpp repo linked above has code + links for gemma.cpp specific weights)
If so, you can find the weights here https://www.kaggle.com/models/google/gemma - each of the "model variations" (flax, jax, pytorch, keras, etc.) has a download for the weights and links to its code.
If you're comfortable with flax, that's DM's own reference implementation: https://github.com/google-deepmind/gemma
Do you have any estimates on getting Metal support similar to how llama.cpp works?
Why `.gguf` files are so giant compared to `.sbs`? Is it just because they use fp32?
You can see the various quantizations here, both for the 2B model and the 7B model. The smallest you can go is the q2_K quantization of the 2B model, which is 1.3GB, but I wouldn't really call that "functional". The q4_0 quantization is 1.7GB, and that would probably be functional.
The size of anything but the model is going to be rounding error compared to how large the models are, in this context.
--
Large language models (LLMs) have achieved significant progress in recent years, with models like GPT-3 and LaMDA demonstrating remarkable abilities in various tasks such as language generation, translation, and question answering.
However, 2b parameter models are a much smaller and simpler type of LLM compared to GPT-3. While they are still capable of impressive performance, they have a limited capacity for knowledge representation and reasoning.
Despite their size, 2b parameter models can be useful in certain scenarios where the specific knowledge encoded in the model is relevant to the task at hand. For example:
- Question answering: 2b parameter models can be used to answer questions by leveraging their ability to generate text that is similar to the question.
- Text summarization: 2b parameter models can be used to generate concise summaries of documents by extracting the most important information.
- Code generation: While not as common, 2b parameter models can be used to generate code snippets based on the knowledge they have learned.
Overall, 2b parameter models are a valuable tool for tasks that require specific knowledge or reasoning capabilities. However, for tasks that involve general language understanding and information retrieval, larger LLMs like GPT-3 may be more suitable.
--
Generated in under 1s from query to full response on together.ai
In practice your priority would be fancy quantization, and just any library that compiles down to an executable (like this, MLC-LLM or llama.cpp)
Pre LLM era (let's say 2020), the hardware used to look decently powerful for most use cases (disks in hundreds of GBs, dozen or two of RAM and quad or hex core processors) but with the advent of LLMs, even disk drives start to look pretty small let alone compute and memory.
Distributing 17GB isn’t a big deal if you shove it into Cloudflare R2.
If you are looking for megabytes, yeah, those "chat" llms are pretty unusable at that size.
As maturity arrives, we'll likely see a handful of competing local models shipped as part of the OS or as redistributable third-party bundles (a la the .NET or Java runtimes) so that individual applications don't all need to be massive.
You'll either need to wait for that or bite the bullet and make something chonky. It's never going to get that small.
---
llama.cpp has integrated gemma support. So you can use llamafile for this. It is a standalone executable that is portable across most popular OSes.
https://github.com/Mozilla-Ocho/llamafile/releases
So, download the executable from the releases page under assets. You want either just main and server and llava. Don't get the huge ones with the model inlined in the file. The executable is about 30MB in size,
https://github.com/Mozilla-Ocho/llamafile/releases/download/...
Was not very impressed with the chat either.
So maybe this is neat for embedded projects, but if it's Gemma only, that would be quite a sticking point for me.
It could have been long context? Or a little bigger, to fill the relative gap in the 13B-30B area? Even if the model itself was mediocre (which you can't know until after training), it would have been more interesting.
One thing I do suspect people are running into is sampling issues. Gemma probably doesn't like llama defaults with its 256K vocab.
Many Chinese llms have a similar "default sampling" issue.
But our testing was done with zero temperature and constrained single token responses, so that shouldnt be an issue.
I can't share the eval, but it's pretty simple: it asks a question about some data, and is restricted to only answer yes/no (based on the output logits and suggested in the prompt). It's called with 0 temperature and only 1 output token, so sampling shouldn't be an issue.
The model itself is very unimpressive and I see no reason to play with it over the worst alternative from Hugging Face. I can only imagine this was released for some bizarre compliance reasons.
FYI using the NUQ (4.5-bit) quantization improves throughput by about 1.4x.
I wonder if there's some analogies to the 80s or 90s in here.
Gemma.cpp is a highly optimized and lightweight system. The performance is pretty incredible on CPU, give it a try =)
They are driven entirely by their own curiosity and a desire to push computers to the limit. Combined with their admirable low-level programming skills, you get a very solid, fun codebase, that they are sharing with the world.
Do these models have the same kind of odd behavior as Gemini?
Also, there's a lot of magic going on behind the scenes with configs stored in gguf/huggingface format models, and the libraries that use them. There are different tokenizers, but they mostly follow the same standards.
However, be aware that there were some quality issues with quantization initially (hopefully they're resolved but i haven't followed too closely): https://twitter.com/ggerganov/status/1760418864418934922
Has this perception changed or pretty much the same?
Not sure what you mean about Gemma considering it’s not a service. You can download the model weights and the inference code is on GitHub. Everything is local!
There are some very interesting efforts in JAX/TPU land like https://github.com/erfanzar/EasyDeL
if you figure out a money making software/service, you're gonna be tied to that model to some significant degree.