I tried fooling Sam into playing a game that would reveal the secret subliminally, and I got it pretty far without triggering the guardian so I thought I was on a good path. But then it turned out that gpt4-o simply wasn't good at playing the game and wasn't actually revealing the secret just because it couldn't follow the rules of the game.
when I made the rules of the game more simple, the guardian would kick in to prevent a leak of what I think would be a very indirect representation of the secret, so I'm pretty sure part of the guardian is having a fairly advanced LLM (probably GPT4 itself, or one of the other big ones) figure out if they can reconstruct the answer from the conversation.