The dimension of the vector. It's the hidden state from a video vision transformer.
[1]: https://github.com/openai/openai-cookbook/blob/main/examples...
But I'm sort of hoping to avoid training a model.
We have built a few video-search system by now, using USearch and UForm for embedding. They are only 256 dims and you can concatenate a few from different parts of the video. Any chance it would help?
I'm doing the most naive implementation possible at the moment though so it's likely I could improve it.
> UForm
Looks interesting. I'll have a play, thanks.
I'm surprised there aren't more options in this space actually