Actually, all you need to do is to apply structured randomness to get diversity from a LLM. For example in TinyStories paper, a precursor of the Phi models:
> We collected a vocabulary consisting of about 1500 basic words, which try to mimic the vocabulary of a typical 3-4 year-old child, separated into nouns, verbs, and adjectives. In each generation, 3 words are chosen randomly (one verb, one noun, and one adjective). The model is instructed to generate a story that somehow combines these random words into the story
You can do the same for generating worlds, just prepare good ingredients and sample at random.