> image capabilities are explicitly out of scope.
Can anyone give me the "overview from 10,000ft" on how these multi-modal models ingest images? Are images tokenized? Are there image embedding models? Auxiliary vision heads?
Can anyone give me the "overview from 10,000ft" on how these multi-modal models ingest images? Are images tokenized? Are there image embedding models? Auxiliary vision heads?
For multi-modal training data, e.g. HTML pages or PDFs, does the training data interleave the image tokens amongst the text tokens in the same document? Slightly limited, as doesn't get juxtaposition to text in complex ways, just linear placement of images.
>https://en.wikipedia.org/wiki/Vision_transformer >https://huggingface.co/docs/transformers/model_doc/vit
It's delightful that this is practically identical to the NLP architecture with only the tiniest adaptive tweak!