I suspect given they seemed to understand LLMs this was a bizarre in the sense of they’re not semantically similar despite being syntactically similar. A sufficiently powerful LLM should be able to distinguish the difference. Probably the prompt classifier isn’t as powerful as the backing LLM.