The methodology behind this is severely flawed. Nothing can be concluded here.
I wrote a reply to this on Twitter, which was liked by several members of OpenAI’s staff (to the extent that counts as confirmation):
> If you don't reset the session before each question these results don't mean much — prior answers are included in the prompt and serve as de facto k-shot examples, forcing later answers to be consistent with whatever opinions were randomly chosen at the beginning. n=4, in effect.