Yes, this is essentially it.
As a corollary, the output classes can be any set, rather than needing to be set before training.
As a corollary, the output classes can be any set, rather than needing to be set before training.
My guess would be option 1. Didn’t read the kev repo here which would also explain
softmax(encode(input)*learned_weights)
You have
softmax(encode(input)*encode(categories))
I'm not sure if Jev does it this way, but it's how you get open-vocabulary zero-shot image classification with models like CLIP [1].