> "Zero-shot learning" is when a model attempts to predict a class it saw zero times in the training data. So, using a model trained on exclusively cats and dogs to then detect raccoons.
Wikipedia on ZSL gives this example:
> For example, given a set of images of animals to be classified, along with auxiliary textual descriptions of what animals look like, an AI which has been trained to recognize horses, but has never seen a zebra, can still recognize a zebra if it also knows that zebras look like striped horses.
Coming back to the blog post with this understanding, this part is interesting:
> The breakthrough in our zero shot object tracking repository is to use generalized CLIP object features, eliminating the need for you to make additional object track annotations
Where "CLIP is a neural network trained on a variety of (image, text) pairs." (From the first source again, so CLIP is a pre-trained model, not the algorithm that you need to make the model.)
Looking a bit at the CLIP repository from openai:
> CLIP matches the performance of the original ResNet50 on ImageNet “zero-shot” without using any of the original 1.28M labeled examples
Here it says they don't use any labeled examples? So it just knows that a fish is called fish? This field is confusing and such a rabbit hole ^^