EmbeddingGemma 2: An open, lightweight multimodal embedding model
blog.google
blog.google
For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.
Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.
If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.
(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)
Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.
If you are only interested in recommendations, clustering, etc and won’t have to do embedding at runtime, that’s it.
There is a lot of basic matrix stuff you can do with embedding vectors that are much easier than their LLM equivalents. I can’t remember the name for the techniques but a major one is translating questions annout a subject eg “Who is Slim Shady?”->”Marshall Mathers”so that your input distribution at runtime matches your offline distribution. Much easier to do with open weight models.
https://developers.google.com/edge/mediapipe/solutions/decis...
I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.
I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing. I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.
(lawyer here) — I’m curious: for what you’re describing, wouldn’t the machine need to be constantly running/updating the embeddings to take updated and new files into account? If so, how would that computation load compare to, say, Spotlight constantly updating its index?
The index you’re thinking of is made once per file.. then you compare and search in embeddings. Add/modify files = asynchronous updating or adding corresponding embeddings using the model in memory onto wherever you persist those embeddings(say SQLite)..
- Text: 78 embeddings/second on small texts
- Image: 4 embeddings/second
- Audio: 6 embeddings/second for 30 second chunks (which is how the encoder works)
- Video: 0.2 embeddings/second per minute of video (not entirely surprising since the encoder does 1fps).
Not bad numbers for a local laptop running a model this large, although the M3 Pro is a few years old and I suspect a M5 Ultra will 5x them at minimum.
Per https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2... the audio data mentions speech, environment sounds, and "acoustic events" and for cross-modality "paired text, image, video, and audio data" so I imagine its audio gets encoded mostly to transcribed language and labels like "bells" without rich structural content.
You might want to actually try giant models with native vision (audio optional) input on spectrograms for audio. Depends on what you need embeddings for and if you're able/willing to extract embedding models out of the LLM (some are more or less amenable) but it's actually insane how good can be at "understanding" audio content's spectral profile.
Try giving opus 5.5 audio spectrograms and ask it to tell you what's inside or even to change the notes/words :)
*summons a minimaxir*
740M total (270M text, 170M vision, 300M audio)
Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.
This may gain more traction if they lead with multimodal input decision making.
but you can use this new one and enable/disable what you don't need.
can keep only text for ex.