How does it work? Does the encoder specifically pass the bounding box coordinates to the rest of the network as part of the embeddings?
Broadly it can be said that LLMs work by reducing uncertainty. Interestingly, human consciousness is also theorized by some to work the same way, react to the input in ways to reduce uncertainty.