If a model were actually capable of scheming, it would also have enough situational awareness from its training corpus to know that <thought> parts are monitored too.
If the monitor catches the model writing "let's deceive the user", it's definitely scheming. But if the monitor finds nothing, you've learned almost nothing.
<absence of evidence != evidence of absence>