So what are some use cases people have found for running these sized multimodals on their device? What is it accurate on, and what is the hallucination rate like?
The image search in the Edge Gallery app of theirs is really incredible. Blows their Screenshot app out of the water for what I wanted it for. I stopped using it when they quietly changed their wording around it suggesting it would use Gemini instead of Gemini Nano, which was the whole reason I was excited about it.
I haven’t tried it yet, but, for law practice, I could imagine using it to search a case file for “undamaged roof before Hurricane Katrina” and “damaged roof after Hurricane Katrina” and being able to locate both deposition testimony and pertinent photographs in the body of evidence.
This is exactly the use case. And if your indexing software uses it correctly you can also do "video deposition with Jane Doe where she angrily discussed John Doe's mother's behavior concerning his grandfather's estate." With audio embeddings, transcripts/captions, and periodic image embeddings contributing to tags on the media item you'd be golden.
The main one is mapping images to text and visa versa, e.g. semantic search of images via text, where the images are encoded and the text question is encoded with the same model, then finding nearest neighbors.
I run offmetaedh.com off of them and this model is a drop in replacement for gemma embedding v1 and HQ-CLIP. Better quality, less RAM usage, faster embeddings, I'm fucking pumped for it.