The biggest upside to using a zero-shot classification model is that it's actually zero-shot, i.e., you don't need to construct task-specific training sets.
If it needs to be fine-tuned, then why not fine-tune smaller and cheaper models that are only as big as the task requires?