Images are tokenized. Rumor and greatest likelihood is that it's a ViT.
>https://en.wikipedia.org/wiki/Vision_transformer >https://huggingface.co/docs/transformers/model_doc/vit
For multi-modal training data, e.g. HTML pages or PDFs, does the training data interleave the image tokens amongst the text tokens in the same document? Slightly limited, as doesn't get juxtaposition to text in complex ways, just linear placement of images.
It's delightful that this is practically identical to the NLP architecture with only the tiniest adaptive tweak!