Thanks, but I think that lucb1e's confusion was probably the same as mine -- given pretrained CLIP features, how is this translated to zero-shot tracking?
Are initial bounding boxes given as usual, or are objects of interest created automagically?
Or are they tracked just from text descriptions?
Lots of questions after reading the post :)