25t/s prompt processing
63t/s token generation
Overall processing time per image is ~15secs, no matter what size the image is. The small 4B has already very decent output, describing different images pretty well.Steps to reproduce:
git clone https://github.com/ggml-org/llama.cpp.git
cmake -B build
cmake --build build --config Release -j 12 --clean-first
# download model and mmproj files...
build/bin/llama-server \
--model gemma-3-4b-it-Q4_K_M.gguf \
--mmproj mmproj-model-f16.gguf
Then open http://127.0.0.1:8080/ for the web interfaceNote: if you are not using -hf, you must include the --mmproj switch or otherwise the web interface gives an error message that multimodal is not supported by the model.
I have used the official ggml-org/gemma-3-4b-it-GGUF quants, I expect the unsloth quants from danielhanchen to be a bit faster.