These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.