Did you validate this by running a A/B test? Main question is were you able to classify back into your known categories correctly all the time, or did the errors compound from the llm hallucination plus embedding search
But no classification is perfect. In search in particular, you will also want to have places for manual intervention for high priority queries.