I have 80,000 voice recordings to classify. Very few are in English, and the classes are not in English. Many of the classes are project names or other proper nouns. How could a system not trained on data specific to the problem possibly be expected to work?
and you might find the scores it gives back about its confidence are useful for escalating to a more expensive classifier.
but also, there's no requirement to use it as a zero-shot classifier, you can provide it as many examples as you like, and prioritise giving it examples it had previously gotten wrong. and with input caching it might be economical.
I'd be surprised if they or others don't start offering a fine tuning API for models like this, like openai does for some models (or used to, I haven't checked in a long time).
As are, honestly, most projects I work on.
There are already Jev-shaped open weight local models coming out that can be fine-tuned on specific data. We'll see which of them the community settles on.