>such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests.
From my POV you're over-focusing on a very specific failure story and neglecting a broader swath of possible failure scenarios.
>Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?
The notion of telling an AI which may not, itself, be aligned to solve the alignment problem seems a little dicey.