MetaCLIP – Meta AI Research
github.com
github.com
My inference function (model.predict("image.png")) return an sv.Classifications object that you can load into supervision for processing (i.e. get top k) [1].
The paper [2] notes the following in terms of performance:
> In Table 4, we observe that MetaCLIP outperforms OpenAI CLIP on ImageNet and average accuracy across 26 tasks, for 3 model scales. With 400 million training data points on ViT-B/32, MetaCLIP outperforms CLIP by +2.1% on ImageNet and by +1.6% on average. On ViT-B/16, MetaCLIP outperforms CLIP by +2.5% on ImageNet and by +1.5% on average. On ViT-L/14, MetaCLIP outperforms CLIP by +0.7% on ImageNet and by +1.4% on average across the 26 tasks.
[1] https://github.com/autodistill/autodistill-metaclip [2] https://arxiv.org/pdf/2309.16671.pdf
- You could predict a class (from a static list such as [dog,cat, ...]) or ...
- You could use image embeddings disconnected from text (you could tell image look-alikes but not what they actually represent). By "embedding" text and images in the same latent space, you can now query your images with text query (such a "a large dog") and find the relevant photos. CLIP understands semantics but also is not limited to a set list of classes (thanks to the ability to use of web data in training).
This is a list compiled by OpenCLIP of high performance models (some better than MetaCLIP) for those interested in using CLIP: https://github.com/mlfoundations/open_clip/blob/main/docs/op...
CLIP is a very interesting development in AI these days, so demystifying it is a great idea. Is anyone using CLIP or similar models daily and will find this research useful -- and willing to discuss it? I'm curious what you're doing.
We are generating embeddings on the camera and send them out via chunk-data on the GigE Vision 2.1 protocol.
I work for a computer vision company. I use CLIP almost every day. Example use cases for which I have used CLIP:
- Image classification
- Automated labeling for classification models
- Image clustering
- Gathering images for model training that are sufficiently dissimilar from existing samples
- Content moderation
CLIP is also being used widely in new research. SAM-CLIP, shared with me by my manager today, is using CLIP (https://arxiv.org/abs/2310.15308) and knowledge distillation for training a new model. I have seen references to CLIP throughout multimodal LLM papers, too, although my knowledge of multimodal model architectures is nascent.Do different models do better or worse at this? Is this just untuned outputs? Like are there parameters I should be tweaking? Sorry I'm not able to give too much more detail. I'm mostly using it within A1111's "Interrogate CLIP" option but I've tried using a model I found on replicate as well as installing locally. Same results every time.
It seems vaguely useful but like it misses the mark a lot of the time. I'm assuming I'm doing something wrong.
LLaVA is a multimodal language model you ask questions of. If you don't provide a question, then the default is "Describe this picture in detail". But if you have a concrete question, you're likely to get better results. You can also specify the output format, which often works.
(Make sure to use --temp 0.1, the default is far too high.)
It runs very slowly on CPU, but will eventually give you an answer. If you have more than about four-five pictures to caption, you probably want to put as many as possible as the layers on the GPU. This requires specific compilation options for CUDA; on an M1/M2 it's possible by default, but still needs to be turned on. (-ngl 9999)
This means the resulting caption is of the form "[BLIP caption], [category1 item], [category2 item], ...". It's very rudimentary.
To clarify: CLIP can tell you if a text label matches an image. It can't generate a caption by itself.
There are more advanced captioning methods, but I'm not sure if they're exposed in A1111 (I haven't used it in some months)
https://github.com/salesforce/BLIP
I built a tiny Python CLI wrapper for it to make it easier to try: https://github.com/simonw/blip-caption
- Image search (aka Google Image) to find relevant photos given a text prompt. This is the biggest use case
- Automated Labeling (to answer questions such as "How many entities have attribute Y"?
We didn't fine tune CLIP though. If folks have done it successfully, I would be love to hear them!
But, I would say that the main issue with CLIP is not performance, but that textual input is limited to 77 characters.
This is a severe limitations, if Meta or other company collected the dataset that allowed model with 1024 characters instead it would enrich the word of open source models much more than +2% accuracy.
My hope is that next person or company who works on that will invest into longer context size for text input :fingers_crossed:
https://play.google.com/store/apps/details?id=com.codylab.se...