> This absolutely is not the first Llama3 vision model. They even quote it's performance compared to Llava.
Although this is true, there have been earlier Llama3 based vision releases, none of the latest Llava releases are Llama3 based.
Although this is true, there have been earlier Llama3 based vision releases, none of the latest Llava releases are Llama3 based.
It is not by the original group who have published a series of models under the Llava name.
Llama 3 outputs text and can only see text, this is a vision model.
>that would make it Llama-2-based.
It's based on Llama 3, Llama 2 has nothing to do with it. They took Llama 3 Instruct and CLIP-ViT-Large-patch14-336, train the projection layer first and then later finetuned the Llama 3 checkpoint and train a LoRA for the ViT.