1. The most obvious: keep going deeper. How many layers until it breaks down.
2. The hidden information variant: Can it do a layer where only Sharon has read the previous dialogs, and she has to explain what she read to Doug, and Doug often asks questions to elaborate on things he doesn't understand?
3. The same characters at multiple layers: Can it make a dialog about Jane and John at a later point in time discussing their own earlier dialog? In other words, can it reliably make the distinction between "you" (the object of discussion) and "you" (the the person you're discussing with) for any value of "you"?
4. The tripartite state: Can it simulate dialog with 3 people? 4 people? how many until it breaks?
5. The infinite meta layer: What happens when you ask it to simulate a dialog between itself and yourself, and as part of that dialog you give it this prompt asking it to simulate this same conversation, causing this conversation to appear as a dialog within itself?
Lastly just to remark, I notice that Mary and David are nearly making the same arguments about Alice and Bob as Alice and Bob were making about John and Jane. The formula for it seems to be to introduce to new characters one layer up, have them each pick a side, then fill in roughly the same arguments again. Maybe this pattern is just spurious, but I'm deeply curious to find out if we have fooled ourselves already with just your example. Do further iterations of "two new characters describe two previous characters" result in the same loop over and over, or will it sometimes generate something novel? I'm deeply curious and don't have GPT4 for myself yet.