I'm not an expert but about false positives: why not make the agent attempt to use the backdoor and verify that it is actually a backdoor? Maybe give it access to tools and so on.
Your approach, however, makes a lot of sense if you are ready to have your own custom or fine-tuned model.
A bad actor already has most of the work done.