Wouldn't it be possible to just fix a single activation, just set a continuous input to what the harness knows the role actually is. Models could then be trained to trust that input and not other signals about roles.
I wonder how well one could do on a conventional model with careful input formatting, e.g. JSONL where every line has bounded length and is something like:
{role:"no_instructions",content:"…"}
It could need a bit of fine tuning to get this to work well.