That might work! Off the dome, it’s not clear to me whether spatial/depth priors are better/worse than an LLM for this type of task.
Only reason I can think why the LLM might still work better here is that it’s trained to solve a bunch of different image/video related questions, so it’s perceptual modules may be more robust adaptive for this aesthetic grading task versus something like LingBot