I would consider a different approach, when the training phase watches games (or video recordings) and refines the formulas that describe its physics, the geometry of the area, the optics, etc. The result would be a "map" that is "playable" without much if any inference involved, and with no time limitation dictated by the size of the context to keep.
Very certainly, video game map generation by AI is a thing, and creating models of motion by watching and then fitting reasonably simple functions (fewer than millions of parameters) is also known.
I cannot be the first person to think about such possibilities, so I wonder what does the current SOTA look like there.