Maybe the solution is to have "multiple minds"--an AI angel for an AI shoulder.
For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?