You can ask GPT-4 for bounding box of each object in json format and it’ll give it to you. Don’t know if there is understanding, but it is definitely useful
Broadly it can be said that LLMs work by reducing uncertainty. Interestingly, human consciousness is also theorized by some to work the same way, react to the input in ways to reduce uncertainty.