I think this is similar to how Gemma 4 12B is implemented, but even then I don't think the single layer image embedding is "aware" of the context.
I think this is similar to how Gemma 4 12B is implemented, but even then I don't think the single layer image embedding is "aware" of the context.
In general an embedding doesn’t have intent or awareness in the way you’re looking for. “Embedding” just means one mathematical structure stuffed inside another. So for example the real number line is embedded into the Cartesian plane as each axis- that’s an embedding.
Now in this case specifically, the embeddings in any kind of transformer model encode the meaning of the thing they represent into vectors (which is what the model itself actually operates on). You can train the embedding to be more useful for a particular task at inference time, which already happens.
So embeddings are used in vision models to convert problems which are about the content and meaning of images into problems which are about multiplying matrices. The model doesn’t want to work with pixels (that’s what very basic vision models do, but it tends to be limited to special purpose applications) it wants to work with concepts in the image. That’s what the embedding gives it.
I still don’t really know what you mean about giving the embedding model some context. It embeds whatever you want to embed. So if you want to give it just a jpeg, fine. If you want to embed a jpeg and a json blob with some additional metadata/“context”/whatever, that’s also fine. That’s already how embeddings work.
I'm not sure how practical it is to train that architecture though or whether there would be performance issues.