They literally lack basic relational understanding so
> knowledge across the whole scene
isn't really correct.
> knowledge across the whole scene
isn't really correct.
Relational understanding is about translating the text to the scene.
Whole-scene knowledge is about making sure both eyes are the same color (for example).