It looks interesting but I don't buy that an AI or any human can convince a sufficiently motivated and capable individual to let them out of the prison with no materialistic profit. (like the two defeated individuals)
Let's assume the gatekeeper is a cold hearted psychopath or a person with AI phobia/paranoia to the extreme.
Why would they let the AI out when they can't feel anything for it?
The author does explain the time gap but what if that is only for collecting information about the person beforehand in order to blackmail them or steer the conversation into a pinching point? What if you start with s person with no identity?
There are no ethical concerns here. Maybe author will shout horrible things enough times and since you as the gate keeper needs to keep talking and engaging, you may let the AI out but well, we have a psychopath here.
Do people need to engage in good faith with AI? Can I continue to say Sorry, I can't answer that.? Yes? Does the gatekeeper need to be honest? Can I use a client side toxicity filter or censor certain words?
There is nothing that would restrict above so what if I censor AI from saying let me out or similar phrases?
Can you increase the handicap for the AI?