I haven't seen any evidence an LLM is trainable to be a decent detector for anything people have made any kind of attempts at trying to get past them. Which is as expected as access to a detector effectively makes the problem equivalent to the halting problem (you can tweak the output using a detector as judge until you have a process to bypass it). Some of them are somewhat able to recognised "raw" output.