The AI dungeon thing is a salient point - GPT2 (which it started out with) could barely make a coherent paragraph, and GPT3 could write about half a page of text that more or less made sense.
With modern LLMs, they still get occasionally tripped up, but you could go for pages without a minor detail not making sense.
Something similar might happen with these game models, given enough time.