The generated images are only vaguely similar in detail to the originals, but the fact that they can estimate the macro structure from audio alone is surprising. I wonder if there's some kind of leakage between the training and test data, e.g. sampling frames from the same videos, because the idea you could get time of day right (dusk in a city) just from audio seems improbable.
EDIT: also minor correction, it's not an LLM it's a diffusion model. EDIT2: my mistake, there is an LLM too!