Forgive my ignorance: are these scenes rendered based on how a scene is expected to be rendered? If so, why would we use this over more direct methods (since I assume this is not faster than direct methods)?
For instance if the scenes are a blob of input weights, what would it look like to add some noise to those, could you get some cool output that wouldn't otherwise be possible?
Would it look interesting if you took two different scene representations and interpolated between them? Etc. etc.
Considering their AI achieved about 96% accuracy to the reference, it would be more interesting to see how Blender does on fitting hardware and with a matching quality setting. Or maybe even a modern game engine.