I took the point as: don't make the LLM the classifier. Use it to turn messy input into useful features, then let a normal model make the actual decision. That gives you thresholds/calibration you can inspect.
What I'm not sure about is how stable those features are when you switch the underlying LLM or model version.