The idea that I'm proposing basically isn't an embedding (which is context independent) but rather combining the embedding model with the context of the LLM. It sounds like the embedding model here is normally a "vision transformer" that maps image chunks into tokens with positional embeddings for both the position in the context as well as position within the image. Maybe it could be given the ability to consume the context and decide to emit multiple tokens for a single chunk? So for example, let's say that I ask a question "How many blades of grass are in this image" (a hard question for traditional image embeddings in models since the embedding won't contain the information). If the proposed architecture was both aware of the context and able to "decide" to emit multiple tokens for an image, then for the above it could emit tokens that represent the answer to the question posed, instead of just being the embedding of the image. Or you could ask something like "How many green pixels are there" and again I think it would work better under this architecture than it would otherwise.
I'm not sure how practical it is to train that architecture though or whether there would be performance issues.