Beating OpenAI CLIP with 100x less data and compute
unum.cloud
unum.cloud
Lot of tricks put together for a great final result it seems
It is probably worth writing a paper about, but we are just too busy building tons of open-source stuff. Check out the GitHub org here: https://github.com/unum-cloud
It is not just about the tranformers, but also about databases, networking, and improving the modern data stack for very large scale retrieval-based AI. A lot of the pieces may be pre-production, but I believe the amazing HN community may still enjoy the ways we use io_uring, SIMD, and a few other less then popular technologies.
Fairly standard for any perf-minded shop but good to see more people discovering them.
For some reason the conference hasn’t made the last years talks public or searchable, but you should be able to access it with a link
In other words, can we use it for commercial purposes for free?
It might be a good time to reread What Color Are Your Bits:
This is also informed by existing computer science objectives surrounding indexing, clustering of data and efficient search over data features.
> Our studies of CLIP in a zero-shot setting show that the model displays significant promise for widely-applicable tasks like image retrieval or search. For example, it can find relevant images in a database given text, or relevant text given an image. Further, the relative ease of steering CLIP toward bespoke applications with little or no additional data or training could unlock a variety of novel applications that are hard for us to envision today, as has occurred with large language models over the past few years.
> We trained on the setup of 3x workstations, with 4x RTX 3090 consumer-grade GPUs in each, connected over 200 GBit InfiniBand HDR.
ok so 85x improvement on the GPU count (i suspect even better once you take into account the differences in consumer grade GPU) but i must still be missing something - where does it say it uses 100x less data?
Let's say you came up with the custom model that gives good results, how do you transfer that model so it can be used in an API?
As a general thing, you'd take a request that would require an inference step, which would then invoke the model with some parameters and input, and return the output. Beyond that, you'd need more detail.
The challenge to support a new model architecture is about coding the preprocessing for inputs (like tokenization or image resizing and color feature extraction) and post processing the outputs (for example entity recognition needs to lookup the entities and align the text).
Once an architecture is coded for the pre/post processing, then serving a new model for inference with that architecture is easy!
If you can figure out pricing primarily based on usage you can capture a whole segment of this market.
We have an source project UKV, that partly overlaps with vector-search: https://github.com/unum-cloud/ukv
Another one - UNSW, is a placeholder for now: https://github.com/unum-cloud/unsw
Both will be soon available on cloud marketplaces, but server-less options are a bit harder to cook. Our Discord is the best place to continue conversation: https://discord.gg/Bbh2bjNhvz
Thank you for advice!
I don't love the way their tables[1] report performance though. My understanding is that the "Dataset" column in the table represents the size of the training dataset, not the size of the dataset they are evaluating on. Note that this undersells their performance though, so it isn't like they are trying to hide something here!
Also I'd love to see someone do a similar benchmark for the OpenAI CPT-3 embeddings. I'm pretty unclear how well they compare to something like FLAN-T5, because they don't seem to be evaluated anywhere in the retrieval setting (unless I've missed it?)
[1] See "Zero-Shot Image Retrieval, English-only" in https://www.unum.cloud/blog/2023-02-20-efficient-multimodali...